2 days with the bleeding edge of AI gave me new insights, but mostly affirmations.
Last week I attended Mistral’s AI Engineer Paris, the type of small-scale event where you can really get into the nitty-gritty with attendees and speakers.
Here are my takeaways:
1. Everyone’s building a software factory and still figuring it out.
When you’re working at an AI company and your mantra is to ship at Kilo Speed, sometimes it’s good to get grounded in reality, touch grass, and hear what others are doing. This was one of those moments for me. Looking at the schedule ahead of time, I knew the AI software factory was going to be the buzz of the moment. And it was.
First, what is this software factory thing?
At its core, it’s a heavily automated development setup where AI agents act as proactive workers across the entire software development life cycle (SDLC), managed by human engineers, for now.
In the traditional, human-driven SDLC you loop through ideas, specs, implementation, and shipping to production. The AI software factory aims to ship features and bug fixes faster, raise code quality, and cut the token cost per feature. In practice, that can mean an agent kicking off when an incident happens, going through logs and metrics, sharing its findings in Slack, and opening a PR for the human responder to review. Or working straight from your task management tool to turn PRDs into PRs. Or taking it up a notch and letting the agent decide what to improve.
The promise is real. AI-native companies keep sharing their results: Stripe built Minions, its one-shot end-to-end coding agent; Ramp’s internal coding agent powers 30% of its engineering pull requests; and OpenAI’s Codex has reportedly “taken over” its engineering.
Then there’s reality
Over the past months the industry has gone from tokenmaxxing and blown annual budgets to a more reasonable approach: company-wide AI policies for writing and software development, and automated agentic work across CI/CD to put a halt to the slop grenades (aka stopping the “meat proxy” we all know, who tosses AI-generated slop over the fence for their colleagues to deal with). @thekitze’s slide summed it up perfectly.
Ah, the beauty of giving everyone unlimited tokens and a mandate to ship 10x more work.
There’s still a lot to figure out. Do you let the agent decide what to build, or keep that with a human? How do you review the work, enforce coding style, and manage context and memory? It’s a long list.
Alpic’s Nikolay Rodionov shared how his team tried to automate the whole SDLC loop, but the agent failed to handle product requirements work. Every skill they tried produced worse thinking than a person working through the problem, and a research agent pulling in competitor reports, user feedback, and analytics was the only thing that helped. Everywhere else, from strategy to QA and rollout, agents added real value to their software factory.
At Kilo we see massive productivity gains across every function of the company, but it takes time and trial and error to build your software factory, all while making sure you can trust all the moving parts inside of it.
2. PRs are the bottleneck for the factory
With everyone generating tons of code, review is the big blocker. I chatted with Matt Pocock and caught his talk on fixing it, and he talked about the factory parts:
Accelerators are the inputs: bug reports, support tickets, slow queries, logs, metrics.
Brakes slow the factory down and raise the quality of what comes out.
The first brake is automated, deterministic checks: linting, tests, typechecking, and code quality, all running before a human looks at the PR. Watch out for structure-sensitive tests that know too much about the internals of the system, and for mocking behavior instead of testing it, because both can make your tests lie. He shared some good skills to help with this.
The second brake is automated review, and it should push commits instead of piling comments onto humans. Then zoom out every week: a retro on every agent PR shows which new checks and coding standards to add, so the next review is easier.
3. Small models decide. Big models write.
Brakes cost tokens too. Every check, review, and guardrail adds another model call or more latency, which puts pressure on the third factory objective: cost per feature.
TypeSafe’s Jev launched two weeks before the conference. It’s a “decision model”: instead of text, it returns a decision with a probability, cheap and near-instant, which makes it a handy part of a factory, for example for natural-language PR linting.
We bet on the same trend before Jev with Auto Efficient, which works out what kind of task a request is and routes it to the cheapest model that’s proven it can do that job on our benchmarks. Cheaper intelligence also tends to get used more. A software factory doesn’t need bigger models everywhere; it needs the context and rules to know which kind of brainpower each step needs.
4. Guardrails are safety brakes
Automated checks catch bad code. They don’t catch an agent pasting an API key into a prompt, or following instructions someone hid in the support ticket it was asked to triage. Those tickets, logs, and bug reports are the accelerators feeding the factory, and they’re also the easiest way in. Agents read input an attacker can control, they may have privileged access, and they don’t behave the same way twice. Mistral’s keynote proposed three layers: harden the model, give the agent only the access it needs (in practice, a sandbox), and run guardrails that block a violation while it happens.
5. Own your factory
Several talks were about sovereignty and owning your intelligence. It plays a role in the software factory too, because if one vendor decides which model you use, where it runs, and what happens to your code, you’re renting your factory. You should use local and open models in your stack to protect your data and prevent getting locked in. I wrote a practical guide for engineering teams working through this.
Where Kilo fits
A software factory is only as trustworthy as its parts, and that’s what we’re building at Kilo, now as part of Anaconda. Anaconda has spent years helping enterprises trust the open-source packages and environments they build on, and Kilo applies the same thinking to the agents writing the code. Agents run in sandboxes, so a mistake stays contained, and runtime guardrails with Enkrypt let an admin set policy once. Safety scores sit on every model card, so you see a model’s risk when you pick it. And because you choose the models and providers, you can swap it out when the risk, the cost, or the ownership question changes.
What I’m taking from it
No one has fully figured out the factory, but the gains are real. The next months will be all about building the factory with parts you can trust, and knowing when to slow down to speed up again: human reviews, deterministic checks, and guardrails for when an agent goes off script.
A factory that ships 10x the code only helps if you can trust what comes out of it.




