thezakulo
out of the noise

Weekly · AI model usage · week 40, 2026

space-bunny-alpha + 6 more AI models by real usage · week 40, 2026

What developers actually ran this week, by real usage — space-bunny-alpha, deepseek-v4.1-flash, glm-5.3-flash, and more: the most used, the biggest climbers, and the new entrants, drawn out of the noise.

Most Used

stealth/space-bunny-alpha

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window. Space

1M ctx · 30 tokens this week

View on OpenRouter →

What we said about space-bunny-alpha

Let's talk about Space Bunny Alpha, because the adoption numbers here are hard to ignore. This anonymous model pulled thirty trillion tokens through OpenRouter this week — that's not a typo, that's trillion with a T. So what are developers actually pointing it at? A lot of it comes down to that million-token context window paired with fast inference. People are feeding it entire codebases, long document sets, multimodal inputs — and iterating without constantly trimming context to fit. The adjustable reasoning effort seems to be a draw too: teams dial it down for quick passes, crank it up when they need the model to actually sit and think. Thirty trillion tokens tells you this isn't a few folks kicking the tires — it's running inside real pipelines, real tools, real daily workflows. Whatever the name situation is, developers are clearly voting with their usage. You can see exactly where it lands against everything else on the full board linked in the description. And if you want next week's rankings, go ahead and subscribe.

deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on

1M ctx · 23 tokens this week

View on OpenRouter →

What we said about deepseek-v4.1-flash

Let's talk about DeepSeek V4.1 Flash, because the number here is hard to ignore — twenty-three trillion tokens moved through it this week. That's not a niche experiment. That's developers leaning on it for real, everyday work. What's driving that volume? A big part of it is that million-token context window. People are feeding it entire codebases, long document sets, sprawling logs — the kind of jobs where you'd normally be stitching things together in chunks. And because it's a sparse mixture-of-experts design, only a slice of the parameters fire per request, which keeps it light enough to run at the scale those token counts imply. So you're seeing it pop up in bulk processing, retrieval-heavy pipelines, and long-session agents — anywhere throughput and context matter more than ceremony. That combination of wide context and efficient activation is clearly resonating with teams who need to keep costs sane while handling a lot. The full board is in the description if you want to see where it lands against the rest. Subscribe, and I'll see you next week.

z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while

1M ctx · 10 tokens this week

View on OpenRouter →

What we said about glm-5.3-flash

So let's talk about GLM 5.3 Flash from Z.ai, because ten trillion tokens in a single week is the kind of number that stops you mid-scroll. That's not a niche experiment — that's developers leaning on this thing hard, day in and day out. What are they actually running it for? Mostly two things. First, coding work that doesn't fit in a tidy little snippet — people are feeding it big chunks of a codebase thanks to that one-million-token context window and letting it reason across the whole thing. Second, long-horizon agent tasks, the multi-step workflows that need to hold context without drifting halfway through. The pull here seems to be efficiency. Teams want something they can keep looping in an agent without watching the clock or the bill, and the usage says that tradeoff is landing for a lot of folks. You'll find the full board ranked in the description below. Have a look, and subscribe so you catch next week's numbers.

Biggest Climber

openai/gpt-6-luna

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic

1M ctx · +455 vs last week

View on OpenRouter →

What we said about gpt-6-luna

So let's talk about GPT-6 Luna, because this one jumped off the board this week — usage is up four hundred and fifty-five percent over last week. That's not a slow climb, that's developers reaching for it in a hurry. And when you look at what they're actually running on it, the pattern makes sense. Luna's the fast, cost-efficient option in the GPT-6 line, and that's exactly where it's showing up — high-volume chat, classification pipelines, lightweight agentic loops where latency really matters. When you're firing thousands of calls an hour, a model that stays snappy and keeps costs sane is a very easy yes. It's also carrying a one-million-token context window, which gives teams room to stuff in whole docs or long histories without stitching things together manually. So that's the story: a surge driven by people doing real, repetitive, production work. The full board's waiting for you in the description — go check where everything else landed, and subscribe so you're here for next week's numbers.

x-ai/grok-4.7

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and

500K ctx · +449 vs last week

View on OpenRouter →

What we said about grok-4.7

So let's talk about Grok 4.7, because the jump here is hard to ignore — usage climbed four hundred and forty-nine percent over last week. That's not a slow drift upward; that's developers reaching for it in a big way. And when you look at what they're actually running it for, the pattern makes sense. A lot of this is long-running software engineering work — the kind of tasks where you hand off something messy and let the model grind through it, checking its own work as it goes. That half-million-token context window gives it a lot of room to hold an entire codebase or a sprawling set of docs in view at once, which is exactly what agentic workflows lean on. So the adoption spike tracks with the use case: people are putting it on jobs that take a while and need to stay coherent across a lot of ground. Whether it sticks around at these numbers, we'll see next week. The full board's in the description — go check it out, and subscribe so you catch next week's.

New On The Board

anthropic/claude-sonnet-5.5

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing

1M ctx · 390 tokens this week

View on OpenRouter →

What we said about claude-sonnet-5.5

Claude Sonnet 5.5 lands high on the board this week, and the usage tells a clear story: 390 billion tokens routed through it in just seven days. That's not casual experimentation — that's developers wiring it into the daily grind. When we look at what people are actually running it for, it skews toward the practical middle of the workload: building out features, chasing down bugs, and generating the kind of working code you ship rather than just admire. It's the Sonnet-class pick for well-scoped everyday work, stepping in as a direct upgrade over Claude Sonnet 5. And with a one-million-token context window, teams are feeding it whole repositories and long threads of context without constantly trimming to fit. A number like 390 billion doesn't come from one viral moment — it comes from a lot of people quietly keeping it in their editor loop, week after week. You'll find the full ranked board in the video description, so go check where it sits against everything else. And subscribe so you catch next week's shifts.

openai/gpt-6.1-sol

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional

1M ctx · 258 tokens this week

View on OpenRouter →

What we said about gpt-6.1-sol

Let's talk about GPT-6.1 Sol, because the usage here is hard to ignore: 258 billion tokens routed through it this week. That's not a trickle of people kicking the tires — that's real, sustained workload. So what are developers actually running it for? Sol sits just below Astra in the GPT-6 lineup, and the adoption pattern reflects that. People are leaning on it for agentic coding loops, computer-use tasks where the model's driving tools and clicking through workflows, and document-heavy professional work. That one-million-token context window is clearly doing a lot of lifting — folks are feeding it entire codebases and long document sets without chopping everything into pieces. My read on why it's climbing: it's the practical middle choice. Capable enough for the serious stuff, but priced and positioned so teams can actually run it at volume. And 258 billion tokens says they are. The full board — every model and where Sol lands — is in the description. Subscribe and I'll see you next week for the new numbers.

Get this every week — free newsletter.