thezakulo
out of the noise

Weekly · AI model usage · week 39, 2026

glm-5.3-flash + 6 more AI models by real usage · week 39, 2026

What developers actually ran this week, by real usage — glm-5.3-flash, deepseek-v4.1-flash, hy4-preview, and more: the most used, the biggest climbers, and the new entrants, drawn out of the noise.

Most Used

z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while

1.3M ctx · 19 tokens this week

View on OpenRouter →

What we said about glm-5.3-flash

So here's one that's quietly stacking up serious volume: GLM 5.3 Flash from Z.ai pulled in nineteen trillion tokens this week. That's not a rounding error — that's developers leaning on it hard. And when you look at what they're actually running it for, the picture makes sense. This is a native multimodal model built for efficient coding and long-horizon agent work, the kind of tasks where you fire off a run and it keeps its footing over a lot of steps. It ships with a 1.3 million token context window, so teams are feeding it entire codebases, long transcripts, sprawling logs — and its hybrid sparse and linear attention setup is designed to hold accurate behavior across all that distance. The word to focus on is "Flash." When you're spinning up agents that loop and iterate, cost and throughput add up fast, and that efficiency is clearly showing up in the numbers. You'll find the full usage board in the description below. Subscribe and I'll see you next week with a fresh cut.

deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on

1M ctx · 18 tokens this week

View on OpenRouter →

What we said about deepseek-v4.1-flash

So let's talk about DeepSeek V4.1 Flash, because eighteen trillion tokens in a single week is the kind of number that tells you something real is happening. That's not curiosity traffic — that's developers who've wired this into pipelines and left it running. What are they actually doing with it? A lot of high-volume, latency-sensitive work. With a sparse mixture-of-experts design activating around eight billion parameters on input, it's the kind of model people reach for when they're processing huge batches and every millisecond of throughput matters. And that million-token context window? Teams are feeding it entire repos, long transcripts, sprawling log files — the stuff that used to mean chunking and stitching. The "Flash" in the name is doing honest work here. When you're running document processing or agent loops at scale, cost-per-token and speed become the whole conversation, and clearly a lot of builders have run the math. The full board's in the description — go see where it landed. And subscribe, so you catch next week's movers.

tencent/hy4-preview

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that

1M ctx · 12 tokens this week

View on OpenRouter →

What we said about hy4-preview

Let's talk about Tencent's Hy4 preview, because the usage here is genuinely eye-catching. Twelve trillion tokens moved through this model in a single week. That's not a rounding error — that's developers leaning on it hard, day after day. So what are they actually running it for? The pattern points squarely at coding agents and complex tool-use workflows. This is a mixture-of-experts setup — 49 billion active parameters out of 770 billion total — which means people are reaching for it when they've got multi-step, agentic tasks that need to stay coherent across a lot of moving pieces. And that million-token context window is a big part of the story. When you're feeding an agent an entire codebase, long tool chains, and a running history, you want room to breathe. Developers clearly found that room here. For a preview model to pull these kinds of numbers tells you people are putting it into real pipelines, not just kicking the tires. The full board's in the description — go check where everything landed. And subscribe so you're here for next week's numbers.

Biggest Climber

meta/muse-spark-1.3

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through

1M ctx · +318 vs last week

View on OpenRouter →

What we said about muse-spark-1.3

Let's talk about Meta's Muse Spark 1.3, because the usage jump this week is hard to ignore — up 318% over last week on the board. That's a big move, and when you look at what people are actually running it for, it starts to make sense. Developers are leaning on this one for long-running agentic setups: multi-agent orchestration, coding workflows that span hours, tasks where you need the model to hold onto context instead of losing the thread halfway through. And with a one-million-token context window, there's a lot of room to keep that state alive across an extended job. So the pattern I'm seeing in the numbers is teams handing it the messy, drawn-out work — the stuff where you're chaining steps and you can't afford a mid-task memory wipe. Whether that 318% holds or settles, it's clearly getting real hands-on time right now. The full board's linked in the description if you want to see where it lands against everything else. And subscribe — I'll have next week's usage numbers ready for you.

qwen/qwen3.8-27b

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be

1M ctx · +108 vs last week

View on OpenRouter →

What we said about qwen3.8-27b

Let's talk about Qwen3.8 27B, because this one's climbing fast — usage more than doubled this week, up a hundred and eight percent. That's a real signal. When a jump like that shows up, it usually means developers found a workflow that just clicks. And with this model, it's the vision-language side pulling people in. Folks are feeding it screenshots, diagrams, UI mockups, and pairing that with coding tasks — read the design, write the component. Others are leaning on that massive one-million-token context window for long-running agent runs, where you don't want to keep re-feeding the same documents over and over. Research workflows, professional pipelines, multi-step reasoning that needs room to breathe — that's the territory here. The flexible thinking mode gives teams a dial to turn depth up or down depending on the job. Whatever's behind that surge, developers are voting with their API calls, and that's the only metric this board tracks. You'll find the full rundown in the description below — and hit subscribe so you're here for next week's numbers.

New On The Board

stealth/space-bunny-alpha

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window. Space

1M ctx · 2 tokens this week

View on OpenRouter →

What we said about space-bunny-alpha

Let's talk about Space Bunny Alpha, because the usage here is genuinely eye-catching: two trillion tokens this week. That's not a niche experiment — that's developers leaning on it hard, day in and day out. So what are they actually running it for? A lot of it looks like heavy code work — the kind of tasks where you're feeding in whole repositories and expecting the model to keep track. That million-token context window makes that practical, and the fast inference means people aren't sitting around waiting between iterations. The native multimodal input is showing up in workflows too, where folks are mixing screenshots, diagrams, and text into a single request. The anonymous branding hasn't slowed anyone down. When a model quietly racks up two trillion tokens, it usually means it's just quietly earning its spot in people's pipelines. You can see exactly where it lands against everything else on the full board in the description below. And if you want next week's usage breakdown, go ahead and subscribe — I'll see you then.

typesafe/jev-1.13

jev-1.13, on OpenRouter

1 tokens this week

View on OpenRouter →

What we said about jev-1.13

One trillion tokens. That's the number sitting next to jev-1.13 on the board this week, and it's the kind of figure that tells you something real is happening in developers' actual workflows — not a demo, not a weekend experiment, but sustained day-to-day usage at serious scale. When a model crosses into trillion-token territory, it usually means it's landed somewhere in the pipeline that runs constantly. Think bulk processing, code assistance, agent loops that fire off request after request, document work that never really stops. That volume doesn't come from people kicking the tires — it comes from teams who've wired it into something they depend on and just left it running. We don't have context-window details listed for jev-1.13 this week, so I'll stick to what the data shows: developers are reaching for it, over and over, and the token count reflects that trust in the plumbing. The full ranked board is in the description if you want to see where it lands against everything else. Subscribe and I'll see you next week with a fresh cut of the numbers.

Get this every week — free newsletter.