For most of this year, the enterprise AI conversation has been dominated by a simple, three-part equation: Cost + Performance + Efficiency. Engineering teams tracked tokens per second, context window sizes, and the dollar cost per million tokens.
Maybe millions changed to billions (and now trillions) of tokens. But the equation was generally the same, and it made sense as long as you calculated performance in a way that fit your business use case, whether that’s coding, data science, marketing, financial modeling, or just general building.
But as a new generation of foundation models arrives, that equation is obsolete.
Today it looks more like this: (Cost + Performance + Efficiency) × Security.
Security is no longer a compliance checkbox. It is the multiplier. A fast, cheap model is worth very little to an enterprise if it leaks data, writes vulnerable code, or behaves unpredictably once it runs unsupervised. The scale makes this urgent: Kilo by Anaconda orchestrates trillions of tokens a month for more than 5 million developers, and each of those tokens passes through a model someone had to choose and trust.
This week something shifted in how model labs are addressing this changing landscape and building for that new equation.
Google’s Calculated Caution with Gemini 4 Argon
The capability of modern frontier models is forcing even the largest labs to slow down. Instead of a global GA launch for Gemini 4 Argon, Google opted for a gradual, phased rollout because of safety and security concerns.
Google has been using the model extensively internally and noted that it’s especially impressive around memory efficiency and large-scale codebase migrations and optimizations. But they are working to improve a number of “critical frontier safeguards” before doing a general release.
When a model can carry out deep, multi-step autonomous reasoning, the damage from a hallucination or unsafe output grows quickly. As the WSJ reported (“Google rolls out new AI model gradually amid safety concerns”), Google’s approach signals that managing deployment risk now outranks winning the news cycle.
It’s also important to note that there are already highly capable models out there from Google DeepMind, like Gemini 3.8 Flash which (and not just because it’s currently 50% off…). Gemini 3.8 Flash outranks GPT-5.5 and Kimi K3 on KiloBench and it’s become a daily driver for many Kilo Coders where security and speed are both essential concerns.
OpenAI’s DevDay Surprise: Shelving Astra for Sol
OpenAI offered an even more dramatic example. Going into DevDay in San Francisco, the industry expected GPT-6.1 Astra.
Instead, OpenAI chose not to release the latest Astra, citing a range of safety concerns. And then, after the news cycle was already attached to the fact that a new model had not been released, OpenAI shipped GPT-6.1 Sol at their annual DevDay event in San Francisco.
Sol and Luna models have been especially popular in Kilo and the new model has already proven to be aligned and guardrailed to balance next-generation performance with stricter safety guarantees.
The message seems to be that raw capability that can’t be secured won’t ship. The coming months will tell if that’s a consistent message or just a one-time blip in the big machine.
Anthropic’s Claude Sonnet 5.5: Efficiency and Safety
Anthropic also did something of a surprise this week with the launch of Claude Sonnet 5.5, following closely on the heels of Opus 5.5.
The new Sonnet demonstrates a key focus on balancing performance with real-world efficiency and enhanced security. It’s the most efficient model yet from Anthropic. Priced at $2 / $10 per million tokens, Sonnet 5.5 runs over 30% faster and reduces per-task costs by up to 30% for most work compared to Sonnet 5.
On safety and alignment, Sonnet 5.5 “scores better than Sonnet 5 across our alignment evaluations” and features “new, stronger safeguards to match its jump in capability”. As part of these stricter controls, higher-risk cybersecurity requests return an immediate refusal. If you’re using that equation, (Cost + Performance + Efficiency) × Security, Claude Sonnet 5.5 is an excellent Security multiplier.
That said, GPT-6 Astra and Opus-5.5, not the new Sonnet, that have been high on Kilo’s Code Review leaderboard lately (along with GLM 5.3 Flash, a remarkably powerful flash model from the impressive GLM-5.3 family).
The Open-Weight Paradox: MiMo v2.6 Pro and GLM 5.3
While proprietary labs slow down to build guardrails, the open-weight ecosystem keeps accelerating, with a different security dilemma.
In our internal benchmarks, MiMo v2.6 Pro and GLM 5.3 show state-of-the-art results in autonomous code review, vulnerability detection, and secure code generation. But enterprise teams still ask hard questions. How were these models trained? What is in the weights? And where, exactly, are requests being served?
Even Anthropic has been comparing GLM-5.3 to Fable for cybersecurity capabilities, finding that “GLM-5.3’s lax safeguards significantly increase the cyber capabilities available to malicious actors.”
You shouldn’t have to pick between those capabilities and your InfoSec team’s requirements. Kilo Enteprise gives you several levels of security controls with no data retention on paid plans; subprocessors, encryption, and access controls; incident response; DPA support; and a live Trust Center.
I would also highly recommend checking out Enkrypt AI’s offering. As an enterprise AI security and governance platform, Enkrypt AI provides inline Sentry guardrails, automated red-teaming, and workspace compliance monitoring to detect vulnerabilities, prevent prompt injections, and block data leakage in real time. Incorporating their security and policy layer ensures your organization can safely deploy autonomous models and leverage next-generation performance without compromising on safety or regulatory compliance.
Plus, Enkrypt’s risk scores are now available directly in the Kilo Leaderboard so you can get a sense for how models perform in the wild. We’re here to help improve the new equation for organizations of every size and shape. Happy building!





LLMs are mathematically incapable of being reliable or secure, without considerable external scaffolding that almost obviates the usefulness of the LLM itself.
AI security is an enormous field tacked on top of general cybersecurity which is also an enormous field. There is also GRC - governance, risk management and compliance.
If corporations don't get a handle on all of them, it won't matter what the developers think.