Step 5 Preview, Claude Haiku 5.5, and Glyph Cluster each give you a reason to experiment. Here’s what they’re built for, why they’re worth trying, and what I’d put in front of them first.
Step 5 Preview
Step 5 Preview is StepFun’s model for agentic work: tasks where a model needs to use tools, work through multiple steps, and keep moving toward a deliverable. Its documented use cases include software engineering, research, and analysis across documents. It supports text, image, and video input, with a million-token context window.1
Why try it? Some coding problems require more context than a single function and an error message. You need to connect behavior across files, understand dependencies, and check whether a proposed change actually solves the problem. Step 5 Preview is worth evaluating on that kind of investigation.
Start with a bug that crosses a few files. Give it the error, explain what should happen, and ask it to trace the issue through your code.
A prompt to try:
Trace why this API request fails, starting at the client and following it through the server. Explain the cause, make the smallest appropriate fix, and run the relevant tests.
I’d pay attention to whether it finds the cause and finishes the fix. There’s a difference between explaining a bug convincingly and leaving you with working code.
Claude Haiku 5.5
Haiku 5.5 is built for work where speed and cost matter, especially tasks you run frequently. Anthropic highlights summarization, extraction, classification, routing, and subagent work. It also supports adaptive thinking, giving it room to reason through a task when needed.2
Why try it? A lot of the work in a coding session consists of smaller jobs: understanding a module, adding coverage, fixing validation, or summarizing what changed. If Haiku handles those well in your project, it could make everyday work faster and less expensive. That’s worth testing before you automatically reach for a larger model.
Try it on something small that you actually need done. My suggested starting point in Kilo is a task with a clear scope and an outcome you can check.
A prompt to try:
Add tests for this validation function. Cover missing values, malformed input, and boundary cases. Follow the existing test conventions and run the tests.
If you get useful tests without a lot of back-and-forth, that’s a reason to keep using it. A model doesn’t need to handle your hardest task to earn a place in your workflow.
Glyph Cluster
Glyph Cluster is a stealth reasoning model for coding and analysis across large inputs. Vercel describes tasks such as reviewing code, debugging failures, comparing implementation options, and working through plans with multiple steps. It supports function calling, so an agent can connect that reasoning to tools.3
Why try it? Before you make a substantial change, there’s value in getting a model to work through the dependencies and tradeoffs. A migration is a good example: the code changes might be straightforward, but preserving authentication, error handling, and existing behavior takes more thought. That’s the kind of task I’d use to see what Glyph Cluster can do.
Give it a migration to think through. Vercel’s announcement includes moving from Express to Next.js route handlers, which is a useful example to borrow.
A prompt to try:
Review these Express routes and plan a migration to Next.js route handlers. Identify changes needed for authentication, middleware, error handling, and tests. Propose an order for the migration and flag decisions that need my input.
A useful plan should pick up on the details of your application: which middleware needs replacing, where behavior could change, and what you’ll need to test.
Glyph Cluster currently accepts text input and doesn’t support structured outputs. Vercel also says its stealth offering may use prompts and responses for training, so check the terms for the route you’re using.
Pick one and try it
You probably already have a task sitting around that fits one of these examples.
Open Kilo, switch models, and give it a shot. See whether it gets you to a useful result—and how much work you have to do along the way.
If you try one, tell us what you gave it and how it went. We’d like to hear about the rough edges, too.


