Weekly · AI model usage · week 41, 2026
deepseek-v4.1-flash + 6 more AI models by real usage · week 41, 2026
What developers actually ran this week, by real usage — deepseek-v4.1-flash, space-bunny-alpha, glm-5.3-flash, and more: the most used, the biggest climbers, and the new entrants, drawn out of the noise.
Most Used
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on
1M ctx · 35 tokens this week
View on OpenRouter →What we said about deepseek-v4.1-flash
Developers are reaching for this one when they've got mountains of context to chew through — think whole repositories, long transcripts, sprawling log files — and they want fast, cheap passes without babysitting the budget. DeepSeek V4.1 Flash is doing a lot of that heavy lifting right now, and the numbers back it up: thirty-five trillion tokens flowed through it this week on OpenRouter. That's a staggering volume, and it tells you something about where it's landing in real workflows. With a million-token context window and a sparse mixture-of-experts setup that only lights up a slice of its parameters per request, teams are using it for bulk summarization, document pipelines, and batch processing where throughput matters more than anything else. It's the kind of model you wire into the background — the one quietly handling the grunt work while you focus on the interesting parts. When a model moves thirty-five trillion tokens in seven days, that's not curiosity. That's developers who've found a tool that fits the job.
space-bunny-alpha, on OpenRouter
22 tokens this week
View on OpenRouter →What we said about space-bunny-alpha
Developers are leaning on this one for high-volume pipeline work — the kind of behind-the-scenes processing where you're pushing requests through at scale and watching the token counter spin. That's space-bunny-alpha, and the adoption story here is honestly hard to look away from. Twenty-two trillion tokens routed through it this week. That's not a handful of teams kicking the tires; that's production traffic, automated workflows, and batch jobs running around the clock. A number that size tells you people have wired this into systems they depend on, not just experiments they're poking at. When you see usage climb into the trillions, it usually means developers have found a reliable fit — something that slots into their stack and quietly keeps doing the job. We don't have a published context window on this one yet, so I'll leave the specs alone. But on raw developer pull, space-bunny-alpha is pulling serious weight this week, and it's earned its spot near the top of the board.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while
1M ctx · 10 tokens this week
View on OpenRouter →What we said about glm-5.3-flash
Developers are wiring this one into agent loops that run for hours — multi-step coding workflows, long document pipelines, the kind of tasks where a model has to hold its place across a sprawling session. That's where GLM 5.3 Flash from Z.ai is landing, and the usage this week tells the story: ten trillion tokens. That's not a number you hit from people kicking the tires. It's production traffic — people leaning on it for efficient coding and those long-horizon agent jobs where keeping cheap, fast throughput actually matters to the bottom line.
The million-token context is clearly part of the draw. When your agent needs to stay coherent across an entire codebase or a stack of reference docs, that headroom stops being a spec sheet bullet and starts being the reason teams pick it. Pair that with its native multimodal side, and you can see why builders are reaching for it when the job runs wide and long rather than short and simple. Ten trillion tokens says plenty of them found it fits.
Biggest Climber
GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional
1M ctx · +505 vs last week
View on OpenRouter →What we said about gpt-6.1-sol
Developers are leaning on this one for agentic coding runs and document-heavy workflows—the kind where you hand off a whole repo or a stack of contracts and let it chew through the context without losing the thread. That model is GPT-6.1 Sol, OpenAI's upgrade to GPT-6 Sol, sitting just below the flagship Astra in the GPT-6 lineup. What's catching my eye this week is the usage curve: it's up a staggering 505 percent versus last week, which is the kind of jump you only see when a tool clicks for people fast. A lot of that seems to come from the million-token context window—room enough for teams building agents that browse, operate a machine, or reason across long professional documents in a single pass. When a model more than quintuples its traffic in seven days, that's not curiosity, that's developers folding it into real pipelines and keeping it there. Worth watching where this lands next week once the early adopters settle into a rhythm.
Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing
1M ctx · +282 vs last week
View on OpenRouter →What we said about claude-sonnet-5.5
Developers are leaning on this one for the everyday grind — building out features, chasing down bugs, and shipping clean changes without a lot of hand-holding. That workhorse is Claude Sonnet 5.5, Anthropic's direct successor to Sonnet 5, and the usage jump this week is hard to ignore: up 282 percent over last week. That kind of climb usually tells you something practical is happening on the ground — people are moving real work onto it, not just kicking the tires. The million-token context window is a big part of the story here, because it lets teams drop in whole repos, long logs, and sprawling docs and keep the model oriented across all of it. For well-scoped, day-to-day tasks, that combination of capacity and reliability is exactly what a lot of developers want in their main driver. When a Sonnet-class upgrade lands and adoption nearly quadruples in a single week, it's worth paying attention to where that momentum is coming from.
New On The Board
Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window. It is built for agentic use cases where response
524K ctx · 828 tokens this week
View on OpenRouter →What we said about solar-mini4
Teams are wiring this one into agents that chew through long documents, orchestrate tool calls, and keep context across sprawling conversations without blowing the budget. That's Upstage's Solar Mini 4, and the usage story here is genuinely interesting. It's a compact mixture-of-experts build — 35 billion parameters on paper, but only about 3 billion active per token, which is the trick that keeps it lightweight and cheap to run at scale. Pair that with a 524K context window and you've got something developers can point at big files, logs, or multi-step workflows and let it run. This week it pulled 828 billion tokens across OpenRouter — a serious number for a small model, and a clear signal that people are deploying it in real pipelines, not just kicking the tires. That kind of volume usually means it's landed in production somewhere, handling the high-frequency, cost-sensitive jobs where every token counts. When a compact model moves that much traffic, it tells you developers found a sweet spot worth returning to.
Step 5 Preview is StepFun's flagship model for agentic work, built on a sparse Mixture-of-Experts architecture (27B active / 600B total parameters). It performs strongly in software engineering and professional
1M ctx · 632 tokens this week
View on OpenRouter →What we said about step-5-preview
Developers are wiring this one into agentic pipelines — the kind of multi-step workflows where an assistant plans, writes code, checks its own output, and keeps going without a human babysitting every turn. That's Step 5 Preview from StepFun, their flagship for agentic work, built on a sparse Mixture-of-Experts setup with 27 billion active parameters drawn from a 600-billion pool. What's catching my eye is where the usage lands: software engineering tasks, long professional documents, and workflows that lean on that million-token context window to hold entire codebases in view. And the number backs it up — 632 billion tokens routed through it this week. That's not a few teams kicking the tires; that's sustained, production-level traffic. The big context window seems to be doing a lot of the heavy lifting here, letting developers feed in sprawling repos and documentation without chopping everything into fragments. When you're building agents that need to reason across a lot of state, that headroom matters, and the adoption we're seeing suggests plenty of people have found a real use for it.