When generating costs nothing, verifying becomes the problem
- Agentic AI
- Explainability
Picture the scene. It already plays out every day.
A generative AI tool produces a thirty-page memo in forty seconds: the synthesis of a supplier file, a compliance review, a purchasing recommendation. The next day, that memo must be defended before a committee, a client, or an auditor.
Generation took less than a minute. How long will it take to verify the result? An hour, sometimes a day.
Production has become almost instantaneous. Trust, for its part, is still paid for at the human hourly rate.
This paradox sums up part of AI’s current problem in the enterprise. The world’s four largest cloud players invested on the order of 400 billion dollars in 2025, and are announcing close to 700 billion for 2026. Meanwhile, an MIT report estimated in the summer of 2025 that 95% of generative AI pilots were not yet producing a measurable return, while S&P Global Market Intelligence noted that 42% of the companies surveyed had abandoned most of their AI initiatives, up from 17% a year earlier.
The power of the models is no longer the main obstacle. Neither is access to information.
The problem begins after generation.
AI has shifted the work
A language model is probabilistic by construction. It produces answers that are plausible, often accurate, sometimes wrong. Yet a factual error, a misread source, or an overlooked contradiction does not spontaneously announce itself to the reader.
As soon as the result actually commits a decision, the serious user therefore adopts a simple strategy: verify.
And often, verify almost everything.
The work has not disappeared. It has shifted.
Where the operator used to spend their time searching, comparing, analyzing, and then writing, they now have to check an answer that has already been produced. Generation spectacularly accelerates part of the process, without necessarily accelerating the whole.
This shift matters all the more because human control takes place in an environment that is already saturated. Microsoft’s study of the workday, published in June 2025 from data covering 31,000 people in 31 countries, counts an average of 117 emails and 153 instant messages per person per day, with an interruption every two minutes.
In that context, producing more information does not mechanically increase decision capacity. It can simply move the bottleneck to the person responsible for sorting, verifying, and owning that information.
An error does not announce its presence
One could imagine reducing this cost by sampling: verify a portion of the results and trust the rest.
That strategy works when the stakes are low. For a draft email or a first exploration, an occasional error is often acceptable.
It becomes much harder to defend when the output commits a decision.
Suppose one answer in twelve contains a significant error. To the eye, nothing necessarily distinguishes that answer from the other eleven. That is precisely the problem: if the error were immediately identifiable, it would not need verification.
In the professional uses that matter, an automated output therefore has value only if it can be linked to its sources, understood, challenged, and revised.
A manager must be able to ask why a recommendation was made. An auditor must be able to retrieve the elements that support it. A new piece of information must be able to change the conclusion without erasing the history that preceded it.
The regulatory framework is, moreover, beginning to formalize this requirement. Article 14 of the European AI regulation imposes effective human oversight for high-risk systems.
But regulation here only makes explicit a more general economic problem: as long as a human remains responsible for the final decision, the question is not only what AI can produce.
You have to know what it costs to be able to trust it.
The product is no longer just the answer
A company does not buy a conversation with its data. It buys a solved problem.
This distinction becomes important as AI moves beyond productivity interfaces to take on the work itself. In “Services: The New Software”, published by Sequoia in March 2026, Julien Bek points out that roughly six dollars are spent on services for every dollar spent on software.
In other words, the software market mostly represents the tools used to do the work. The services market represents the work itself.
AI is now beginning to address the latter.
But producing the work is not enough. For an automated result to actually enter a professional process, a human must still be able to use it without having to redo it.
That is where the value is gradually shifting.
Content generation is becoming commoditized. A growing number of models can summarize a document, draft an analysis, or propose a recommendation.
What remains rare is the ability to know why a conclusion is credible, what it rests on, what contradicts it, and what is still missing before deciding.
The next challenge is therefore not simply to produce more.
It is to make trust as scalable as generation.
Preparing the decision rather than making it
Our starting point is deliberately simple: the human decision-maker remains responsible for their decision.
Our role is not to replace their judgment. It is to reduce the cost required to exercise it properly.
For a given class of decision, we first define what a professional should reasonably gather in order to rule: necessary sources, business criteria, contradictory elements, acceptable level of uncertainty, control rules.
Then, for each case, our agents build a case file.
The principle is close to that of a department that prepares a file before submitting it to a validating authority. The authority remains human. The system prepares the elements needed for its ruling.
The file first presents a conclusion and, where relevant, a recommendation. It then lays out a compact chain of evidence in which every element keeps its source and its date. It flags contradictions rather than resolving them silently, makes missing information visible, estimates its confidence level, and lets the detail be unfolded only when it becomes necessary.
The operator no longer has to reconstruct the entire reasoning themselves. They examine it, challenge it if necessary, then rule: approve, send back for completion, or dismiss.
Their responsibility does not disappear.
What decreases is the cost required to exercise it.
This distinction is essential. The goal is not to remove the human from the loop, but to sharply increase the quantity and complexity of information they can reasonably supervise.
An AI that generates a thousand pages per minute is not particularly useful if the human capacity to validate them remains unchanged.
A defensible result requires a defensible memory
Such a file cannot be produced from a memory that keeps only the last known version of a fact.
Classical information systems often seek to maintain a consistent state: one value per subject, updated as new information arrives. When two sources diverge, one usually ends up replacing the other, or a reconciled value is kept.
That logic is suited to knowing the current state of a system.
It is far less suited to explaining how a conclusion was reached.
The real informational environment is rarely perfectly consistent. A source ages. Two credible actors contradict each other. A data point that was correct yesterday becomes wrong today. A reasonable hypothesis is invalidated by new evidence.
To be able to defend a decision, the system must therefore keep more than a final state.
It must keep its history.
To that end, we represent knowledge not as definitively established facts, but as beliefs: sourced, dated, revisable, and associated with a confidence level.
Two contradictory statements can coexist as long as the available elements do not allow them to be settled. Every change remains linked to the evidence that triggered it. An immutable, replayable event journal makes it possible to reconstruct the state of the system at a given moment and to understand how it evolved.
This architecture is therefore not a theoretical refinement added to the product. It is the necessary condition for the result to be audited, challenged, and then updated without losing the trace of the initial reasoning.
Two prototypes run on this foundation today. Our first use cases cover strategic monitoring and supply chain surveillance, in particular from component bills of materials.
Starting where trust is worth the most
This requirement obviously does not have the same economic value everywhere.
To generate a first idea, summarize a meeting, or draft a letter, a merely plausible answer may be enough.
We therefore start where the decision truly commits: defense, the nuclear industry, regulated finance, public administrations, and, more generally, environments in which a conclusion must remain justifiable several weeks or several years after it was reached.
In these sectors, the cost of control is already high. An error engages an identifiable responsibility. Traceability, compliance, and auditability are not side features: they are part of the product being bought.
These environments also impose a demanding conception of sovereignty. Not just local hosting, but the ability to operate without extra-European legal exposure and, where necessary, with models run locally.
We accept the possible performance trade-off that this implies.
Our value does not lie in a few extra points on a model benchmark. It lies in the layer of trust built around that model: memory, evidence, contradictions, auditability, and human oversight.
Measuring the time to trust
One simple question remains: how do we know whether any of this works?
Not by measuring the number of tokens generated. Not by timing the production of a report. Not by showing that an agent can run a spectacular demonstration on its own.
The right unit of measurement is the time it takes for a human to be able to own the result.
We therefore track a simple ratio:
time to review the file / time needed to carry out the same casework manually.
As long as that ratio is not clearly below 1, AI has not really increased the productivity of the process. It has simply shifted the work.
Our ambition is the opposite: that a result produced in a few seconds can be understood, verified, and carried by a human in a few minutes, evidence in hand.
Generation has already become abundant.
The next step is to make trust abundant too.
Sources
- CNBC, “Tech AI spending may approach $700 billion this year, but the blow to cash raises red flags”, February 6, 2026.
- MIT NANDA, “The GenAI Divide: State of AI in Business 2025”, July 2025.
- S&P Global Market Intelligence, via CIO Dive, “AI project failure rates are on the rise”, 2025.
- Microsoft WorkLab, “Breaking down the infinite workday”, June 2025.
- Julien Bek, Sequoia Capital, “Services: The New Software”, March 2026.
- Regulation (EU) 2024/1689 on artificial intelligence, Article 14, human oversight, applicable to high-risk systems since August 2026.