Trending Topics

Ilia Razvin, Founder and CEO of IOSYA: “The future is not an improved yesterday.”
For Ilia Razvin, the problems with AI agents become clearer when viewed through the lens of engineering. The founder and CEO of IOSYA began his career working with electrochemical sensors and process models in chemical environments where a bad assumption could have physical consequences. That experience shaped a straightforward principle: a result without its conditions is not a result. It is a principle he now applies to AI, arguing that an agent that produces a convincing answer is of limited value if it cannot show which source it used, which version it relied on, what decisions were made along the way or when it should stop.
In this interview, Razvin explains why so many AI agents “pass the demo and die in week two”, and why the difficult work of making them reliable happens largely outside the model itself. He discusses how IOSYA handles messy technical documentation, why smaller models can outperform frontier systems on routine production tasks, the practical reality of prompt injection and retrieval poisoning, and where human oversight remains essential. His argument is less about making AI appear more autonomous and more about building systems that are traceable, verifiable and capable of admitting when they do not know.
You’ve said many AI agents “pass the demo and die in week two.” What typically goes wrong once an agent is exposed to a real enterprise environment, and what separates an agent that survives from one that doesn’t?
A demo runs on a curated slice. Three clean documents, one happy path, and a person in the room who already knows what the right answer looks like. Everything in that room is quietly working to make the machine look competent.
Week two, it meets the actual drive. Four versions of the same specification, one of them a scan of a print of a revision. A file called FINAL_v2_use_this. A shared mailbox where the real decision was made in a reply nobody filed. The person who knew which document governs is on holiday. Nothing there is hostile. It is just normal, and normal is the hard case.
Then it fails in four ways, in this order.
It remembers the conversation and not the work. The session closes and everything the company learned in it evaporates. Next week somebody asks the same question and pays for the same answer again.
It cannot show its reasoning. It may well be right, but there is no source, no version, no trace of which document it read and which it ignored, so nobody will sign their name under the output. An unsignable answer is not an answer. It is an opinion with a subscription fee.
It has no plan for its own failure. It does not stop when it should. It guesses, fluently, and fluency is exactly what makes the guess dangerous. In a demo that reads as confidence. In production that is an incident with a date on it.
And nobody owns it. The demo had an owner standing next to it. Week two, the owner went back to their real job.
What survives is dull and structural. Not a general assistant pointed at a company, but a defined process: what starts it, which data it may touch, which actions it may take, where a human signs, what it produces, what it does when it cannot tell. Plus a full execution trace, kept outside the model, so six weeks later you can reconstruct why it did what it did. That is what we build IOSYA around, and it is not the part that photographs well.
My background is chemistry. On a factory floor the last number explains nothing if someone replaced a sensor before it. The demo tests whether the model is clever. Week two tests whether your company ever wrote anything down.
Industrial environments involve hundreds of pages of technical specifications, CAD drawings, tables. How are you solving the problem of getting agents to reliably understand this kind of messy, highly structured technical data?
By not asking the model to do it.
The instinct is to hand over a 300-page PDF and hope. That fails for a reason unrelated to model quality: a document is not a corpus, it is evidence. Evidence has conditions attached, and flattening it into text destroys precisely the part that makes it usable.
So the work happens before the model is involved.
We parse to structure, not to prose. A table stays a table. A merged cell is information — it usually means “this applies to the three rows below.” Flatten it and you have invented a fact. Same for units, tolerances, footnotes, and the small asterisk that says “except for variant B.”
Every fragment keeps its address. Document identifier, revision, issue date, page, table, clause. A chunk that has lost its address is worthless to us however relevant it looks. A measurement without its conditions is not a result. A specification line without its revision is not a specification.
Drawings we handle honestly. We do not pretend a model reads geometry the way an engineer does. We take what is genuinely machine-readable — title block, revision index, parts list, annotations, dimension callouts where they exist — and treat the geometry itself as something a human inspects. Anything else is a demo trick, and it is how you get a confident answer about a part that was redesigned three years ago.
The first question is always which document governs. This is the step people skip and it causes most of the damage. Given four copies of a specification, the system’s first job is not to answer your question. It is to establish which one is current, flag the conflicts between them, and refuse to proceed when it cannot tell. Aviation and law version their documents properly because a document without its history carries no weight as evidence.
And refusal is a feature we build deliberately. “I do not have the governing revision for this section” is a correct output. It is also an output no demo has ever contained.
The result is boring to look at, which is the point. The model does the small final step. The reliability lives underneath, where the sources, versions and decisions are kept in a form the company owns rather than rents.
There’s a tendency to assume the most capable frontier model is automatically the best choice. When have you found a smaller or cheaper model performs better, and when do cost and latency become more important than raw capability?
This is a very complex question, and the industry answers it badly, because the frontier model is the one that photographs well.
Start from the task. A frontier model is for judgment under ambiguity: the question is underspecified, the answer will be read by a human who will act on it, getting it wrong is expensive, and getting it right means weighing things nobody ever wrote down.
Most production work is not that. Most production work is routine. Extract eleven fields from an invoice. Classify an incoming document. Decide which of six routes this ticket takes. Check a value against a rule. These have correct answers.
Extraction against a fixed schema is where the small model wins first. You are not asking for intelligence, you are asking for discipline. A small model with a tight schema and a validator produces less variance than a large model improvising, and variance is the thing that wakes people up at night.
Classification and routing, the same, once you have an evaluation set. Build the evals, measure both, and the argument ends. Nobody argues with a number.
Then latency, which compounds in a way nobody models in advance. A person waits eight seconds once and shrugs. Put fourteen steps in sequence and that is two minutes, and at two minutes the person closes the tab and does it by hand. Your project is dead and no bug was ever filed.
The crossover point is precise: when the task has a verifiable correct answer and you have evals proving the small model reaches it. Capability above that threshold buys nothing, and you pay for it on every call, forever. Keep the expensive model for the steps that are irreversible, ambiguous, one-shot, or read by a human who will act.
There is a catch, and it is why I care about this question at all. Routing different steps to different models only works if the state lives outside the models. Memory, sources, decisions, the trace. If those sit inside one vendor’s session object you cannot move a single step without losing the thread. The model is strong, but replaceable — and that is only true if you built it to be.
Prompt injection and RAG poisoning are often discussed as theoretical risks, but you’ve dealt with them in live systems. What did those attacks or failures look like in practice, and what did you change?
The first thing to say is that they almost never look like an attack. That expectation is the problem. People picture a hacker. What you find is stationery.
Instructions inside a document you were supposed to read: a supplier PDF, a template, an email footer, a comment in a spreadsheet cell, text in a colour nobody renders. Written for a human, or written for nothing at all, and the system obeys because it has no structural way to tell “content I was asked to read” from “instruction I was given.”
Then retrieval poisoning that nobody committed. An obsolete revision still sitting in the folder, indexed with exactly the same confidence as the current one. No attacker in this story at all. The effect on the output is identical to one. This is the version that actually costs companies money and the version nobody gives conference talks about.
And duplicates that outvote the truth. The same wrong figure appears in six places because six people copied a slide. The correct figure appears once, in the source. Retrieval is a popularity contest unless you build it otherwise.
What we changed in IOSYA is structural, not cosmetic.
Retrieved content is data and can never become instruction. Not by asking the model nicely to ignore it — by separating the two at the architecture level, so instruction arrives on one channel and material on another.
Permissions attach to the operator, not to the prompt. An injected instruction can request whatever it likes. If “send external email” is not in that operator’s toolbox, the request has nowhere to go. You cannot talk your way into a capability that was never granted.
Provenance on every retrieved fragment, with untrusted origins marked as untrusted, because an external document is evidence from an unverified source and that is simply what it is.
Write actions get checkpoints. Irreversible actions get a human. Reading is cheap to get wrong. Sending is not.
And a full execution trace: which document, which version, which tool, which decision, in what order. Most teams cannot answer “why did it do that” about last Tuesday. Until you can answer that, you do not have a security posture. You have a mood.
None of this is exotic. It is the discipline a regulated laboratory applies to sample handling, moved into a different building.
When a wrong answer can have physical or financial consequences, how do you decide what the agent is allowed to do autonomously and where a human stays in the loop?
We do not set autonomy as a level. We set it per action, and it runs on four questions.
Can it be undone, how fast, and who pays for the undoing? A draft is reversible in seconds. An email to a client is reversible in a week of relationship repair. A payment instruction is not reversible at all. Three different regimes, and treating them as one setting is how accidents happen.
How far does it reach? One document, one department, one supplier, or everyone.
Can correctness be checked by a machine? If a validator, a rule or a test can confirm the output, autonomy is defensible. If correctness needs a person who knows the site, it is not.
Will someone put their name on it? If the answer leaves the building or enters a record of record, a human signs. Not because the machine is untrustworthy. Because accountability has to land on a person, and no architecture has ever solved that.
In practice: reading, comparing, extracting, drafting, flagging contradictions — autonomous. Anything that moves money, alters a system of record, leaves the organisation, or instructs a human to do something physical — checkpoint.
Then the part people get wrong after doing everything else right. The checkpoint has to be cheap. If approving takes twenty seconds and shows five things — what it is about to do, why, which sources, what changes, how to undo it — people read it. If it shows a wall of text, they approve blind within a week, and you have built a rubber stamp with an audit log. That is worse than no checkpoint, because now you have documentation proving a human agreed.
And the stop conditions are written down before anything runs. What happens when data is missing. When two sources disagree. When confidence is low. How many retries, then what, then who gets told. A system that is not designed to fail will fail anyway, just without telling anyone.
You’ve compared agentic coding to working with “a senior engineer with dementia.” What does that analogy reveal, and what would need to change before you would trust an agent with genuinely critical engineering work end to end?
It reveals that capability and continuity are two separate things, and we have been buying one while assuming the other came with it.
The senior part is real. Per unit of work it is genuinely good. It reads unfamiliar code quickly, it knows patterns, it is often faster than I am. I use it daily and I am not performing scepticism here.
The other part is also real. It will re-derive a decision you already made and land on the opposite one. It will reintroduce, with total confidence, the bug you removed on Tuesday, because the reason you removed it was discussed in a meeting and lives in no file it can see. It contradicts yesterday without noticing yesterday happened. So you end up doing a job that did not exist three years ago: you are the continuity. You hold the thread, and the thread is the expensive part.
The analogy is unfair in one direction and I should say so. A person with that condition knows something is wrong. The machine does not. It has no signal for the edge of its own knowledge, so the failure mode is confident continuation rather than a pause. Confident continuation is the most costly property in production software.
Before end-to-end on anything critical, five things.
Memory of the work, structured and outside the model. Decisions and the reasons for them, rejected options and why, source versions. Not chat history. Chat history is a transcript of a conversation; what you need is the state of a project. That distinction is more or less why IOSYA exists.
Calibrated uncertainty with a real stop — the ability to say “I cannot determine this” and halt, rather than produce something shaped like an answer.
Verification the machine performs, not the human. This is how critical engineering has always worked. We do not trust a bridge because the engineer was clever. We trust it because it was checked, by procedures, against standards, by people who did not design it.
A trace I can audit afterwards. Not out of suspicion. It is the normal standard for any human engineer working on something that matters, and there is no argument for a lower one here.
And a name at the end. Someone accountable. No model release fixes that, and I do not think it should.
Give me those five and “end to end” stops being a frightening phrase, because it will mean roughly what it means in aviation: automated, monitored, traceable, stoppable.
We are not there. We are somewhere useful, which is a different place. The future is not an improved yesterday.
