Data Intelligence

Query massive datasets — public or your own,
and put them to work for your business.

100+ projects delivered
15+ products live
50k+ people trained
Trusted by
UPVUniversity of OviedoIgnisUniversity of MálagaEOISaicaPublic University of NavarreIngeropGroupe ADFEurus EuropeHM HospitalesFenie EnergíaLinkedIn LearningIRYCISCéltica EnergíaCogen EnergíaConnectifInvesydeArticaeQuibusWorkSmart

The problem

From raw data to AI-ready infrastructure

Large datasets almost always start as a mess: daily exports, spreadsheets with dozens of tabs, codes with no context, numbers that never quite match across systems. The data exists, but nothing can be asked of it until someone on the team prepares it first.

Putting AI in front of that does not remove the problem, it amplifies it. The model answers confidently over data nobody reconciled, and the team stops trusting what comes back. The missing layer is not the model — it is the ground it stands on.

How we work

Three steps, always the same

The data exists, but nothing can be asked of it until someone prepares it first. This is the part we do before any AI touches it.

Normalise. Unified schemas, codes cross-referenced against their catalogues, deduplicated records. The data stops depending on who exported it.

Enrich. Sources are joined to each other and given the context they lack, so a question does not require knowing in advance which table holds the answer.

Expose as an MCP server. Ask in plain language from Claude, ChatGPT or VS Code — read-only, with the exact source cited on every answer.

Proof

What this looks like in production

The same thing we build on a client’s data we have taken all the way to a product: the entire Spanish electricity market, open and queryable by anyone.

As a product

The full history of the Spanish electricity market since 2014, already joined and updated daily from official sources, with a free tier and no code to write.

1.3B+ records
2014 history since
Daily refresh
Free entry tier
See Joltio →

Two ways in

On your data, or on the datasets we already run

On your data

project

We build the same system on your dataset: normalisation, enrichment and an MCP server inside your perimeter, read-only and running with the permissions of whoever is asking.

Already built

try it today

Two public datasets already normalised and queryable, with no project in between.

  • Joltio Spanish electricity market
  • Citenza Legislation of Spain, Chile, Italy and the EU

Questions

What people ask before starting

How large can the dataset be?

We routinely work with billions of rows: Joltio holds over 1.3 billion and refreshes daily. The practical limit is rarely volume — it is how many different sources have to be reconciled.

Does our data leave our systems?

No. The server is deployed inside your perimeter and runs read-only. Every query executes with the permissions of the person asking, never as a super-user.

How long does it take?

It depends on the state of the data, not its size. The first thing we do is look at a real sample and tell you how much cleaning stands in the way before committing to anything.

Does it work with Claude, ChatGPT and Copilot?

Yes. Because it is exposed as an MCP server, the same system answers identically from Claude, ChatGPT, Copilot, Cursor or VS Code, without building one integration per client.

What happens when a source changes format?

That is why this is maintained, not just delivered. On the datasets we operate, the change is absorbed before it reaches you — it is the difference between a living integration and a scraper that stops working one morning.

Where we apply it

The sectors where this layer changes the most

Do you have a dataset nobody can ask questions of?