Weekly · open-source AI · week 40, 2026
magpie + 13 more trending open-source AI repos · week 40, 2026
This week's trending open-source AI on GitHub — magpie, short-video-generator-AI, awesome-jev, and more: the newest, Hacker News talk, a subfield spotlight, and the fastest-rising, drawn out of the noise.
Newest this week
Every agent's model. One place. Codex on DeepSeek, Claude Code on Kimi, from the menu bar.
Go · 1,357 stars already
View on GitHub →What we said about magpie
Here's a problem you've probably run into if you juggle coding agents: each tool is basically married to its own model backend. Codex wants OpenAI, Claude Code wants Anthropic, and switching means digging through config files or environment variables every single time. Magpie, from yetone — the same developer behind avante.nvim — sits in your menu bar and quietly decouples the two. You want Codex running on DeepSeek to save money? Claude Code pointed at Kimi for that longer context window? A couple of clicks and you're there.
What I appreciate is the pragmatism. It's written in Go, so it's a lightweight native binary rather than another Electron memory hog, and the menu-bar approach means you're not context-switching out of your editor. Compared to rolling your own proxy with LiteLLM or a manual gateway, this trades some flexibility for genuinely low friction.
If you're cost-conscious or just curious how different models handle the same agent, it's worth a weekend try. And if repos like this are your thing, the newsletter rounds them up weekly — link's in the description.
AI video processing pipeline for generating vertical shorts using LLMs, Whisper transcription, highlight detection and automated editing
Python · 726 stars already
View on GitHub →What we said about short-video-generator-AI
So here's a project that quietly stitches together a bunch of things you've probably been duct-taping yourself. short-video-generator-AI is a full pipeline in Python: it runs Whisper to transcribe your long-form footage, uses an LLM to figure out which moments actually deserve to be clipped, then handles the vertical crop and editing automatically. What I like is that it treats highlight detection as a real step rather than just chopping at fixed intervals — that's the part most homegrown scripts get wrong. Now, is it going to replace an Opus Clip or a Descript? Not on polish. But those are subscriptions with limits, and this is yours to run, fork, and wire into your own workflow. If you're a creator drowning in raw recordings, or a dev who wants to build a clipping tool without starting from zero, this is a genuinely solid foundation. 726 stars this early tells me the demand is real. If breakdowns like this are useful to you, the newsletter link is down in the description — go grab it.
1207 public resources for Jev, TypeSafe AI's System One decision model, indexed by decision pattern. Source citations, dated link checks and scheduled call-site text checks; runtime and performance are not independently tested here. EN
Python · 581 stars already
View on GitHub →What we said about awesome-jev
So here's something a little different from your usual code drop. Awesome-jev isn't a library you install — it's a curated index. Over twelve hundred resources for Jev, TypeSafe AI's System One decision model, and they're organized by decision pattern rather than just dumped into a flat README. That organizing choice is the interesting part. Most "awesome" lists rot within months because nobody checks the links. This one ships dated link checks and scheduled call-site text verification, so you can actually trust that what you're clicking still exists and still says what it claimed. Worth flagging the maintainer's own honesty, though: runtime and performance aren't independently tested here, so treat this as a reading map, not a benchmark. Who's it for? Anyone evaluating Jev for real decision workflows who wants source citations instead of blog-post hearsay. Compared to the typical awesome-list, the freshness tooling is what earns those stars. If you like when we dig into the maintenance discipline behind a repo and not just the star count, the newsletter goes deeper every week — link's in the description.
Talk of Hacker News
Review behavior, not just diffs. Jev prioritizes human attention; OpenAI explains the changes. Local CLI + agent skill + GitHub extension.
JavaScript · 47 points on HN
View on GitHub →What we said about jev-code-reviewer
So here's what caught my eye about jev-code-reviewer this week. Most AI review tools do the same thing — they read your diff, they flag the obvious stuff, they leave a comment. Jev flips the framing. Instead of asking "what changed in these lines," it's trying to reason about behavior — what does this code actually do differently now, and where should a human actually spend their limited attention. That distinction matters more than it sounds. A lot of review fatigue comes from tools that treat a whitespace tweak and a logic change with the same urgency.
Practically, you get it three ways: a local CLI if you live in the terminal, an agent skill if you're wiring it into a larger workflow, and a GitHub extension for the PR crowd. That flexibility is smart — it meets teams where they already work rather than forcing one surface.
It's early, forty-seven points on Hacker News, so temper expectations. But the "prioritize human attention" thesis is worth watching.
If you want more finds like this each week, the newsletter link is in the description.
A langgraph based workflow with a C++ CUDA harness to optimize CUDA kernels
Python · 37 points on HN
View on GitHub →What we said about agentic-cuda-optimizer
Here's one that scratches a very specific itch. Bertaye's agentic-cuda-optimizer wires up a LangGraph workflow to a C++ CUDA harness, and the loop is the interesting part: an agent proposes kernel changes, the harness actually compiles and benchmarks them, and the results feed back in. So instead of an LLM confidently hallucinating a "faster" kernel, you've got measured numbers grounding every iteration.
Why does that matter? CUDA optimization is notoriously fiddly — occupancy, memory coalescing, warp divergence — and it's the kind of tuning where empirical feedback beats intuition. Pairing an agent with a real benchmark closes that gap. Compared to just pasting kernels into a chat model, this is a proper experimental loop, and it's lighter-weight than heavyweight autotuning frameworks if you want to peek under the hood.
Who's it for? CUDA engineers, ML infra folks, and anyone curious how agentic loops perform on genuinely hard, verifiable problems. It's early, but the architecture is a clean template worth studying.
If you like these deep cuts before they blow up, the newsletter link's in the description — go grab it.
Subfield spotlight
An open-source Python project on GitHub.
Python · 126,410 stars
View on GitHub →What we said about MoneyPrinterTurbo
MoneyPrinterTurbo is one of those repos where the star count tells a story about what people actually want to build right now. At its core, it's a Python pipeline that takes a topic or a script and stitches together a short-form video — pulling stock footage, generating voiceover, timing captions, and exporting something you could drop straight onto TikTok or Shorts. What makes it interesting isn't any single piece; it's that it wires together the boring plumbing between an LLM, a text-to-speech engine, and a video editor so you don't have to.
Compared to paid tools like Pictory or InVideo, you trade polish for control and zero per-video cost — and since it's local-ish, your workflow isn't hostage to someone's pricing page. It's aimed at creators experimenting at volume, and devs who want a hackable base rather than a black box.
Just go in clear-eyed: automated content at scale is a quality and ethics question, not just a technical one.
If you like these breakdowns, the newsletter goes deeper — link's in the description.
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
Python · 76,892 stars
View on GitHub →What we said about unsloth
Unsloth has quietly become one of those tools people reach for when fine-tuning starts eating their GPU budget alive. At its core, it's about making training and running local models dramatically more memory-efficient — the kind of optimization work that lets you fine-tune on a single consumer card instead of renting a cluster. What's interesting about this release is the reach: GGUF and MLX support means it plays nicely whether you're on a Linux box with an NVIDIA card or a Mac with Apple silicon, and the model coverage spans the newer Qwen, DeepSeek, and Gemma families plus diffusion work with FLUX.
Compared to something like Axolotl or raw Hugging Face training loops, Unsloth's pitch has always been speed and lower VRAM without you rewriting your whole pipeline. If you're an indie dev or a small team experimenting with custom models, that's the difference between shipping and shelving an idea.
If you want more repos like this landing in your inbox before they blow up, the newsletter link is in the description. Go grab it.
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
Python · 62,220 stars
View on GitHub →What we said about GPT-SoVITS
Alright, let's talk about GPT-SoVITS, because this one's been quietly earning its sixty-two thousand stars. The pitch is simple: give it about a minute of clean audio, and it'll train a text-to-speech model that actually sounds like the person you sampled. Few-shot voice cloning that doesn't demand hours of studio recordings.
What makes it stand out is the hybrid approach — it pairs a GPT-style language model for the prosody and timing with SoVITS handling the actual voice character. That combination is why the output feels less robotic than a lot of the older open-source TTS you've wrestled with. And crucially, it supports cross-lingual synthesis, so you can clone a voice in one language and have it speak another.
Who's this for? Indie devs building accessibility tools, game audio prototyping, dubbing workflows. Compared to something like Coqui or Bark, this leans harder into the low-data cloning use case.
Just — obvious reminder — get consent for any voice you clone. That part's on us.
If you want more finds like this every week, the newsletter link is in the description.
Fastest-rising
Fastest and cheapest web agent
Python · +11,274 stars this week
View on GitHub →What we said about jev-ultrafast
Let's talk about jev-ultrafast, because the name is doing some heavy lifting and, surprisingly, mostly earning it. This is a web agent from the browser-use folks, and the pitch is speed and cost — two things that quietly kill most agent projects the moment you try to run them at any real scale. Web agents are notorious for burning tokens by feeding the model giant DOM dumps on every single step. The interesting move here is optimizing that loop so you're not paying for a novel's worth of HTML just to click one button.
Where this matters is the boring, high-volume stuff: scraping behind logins, filling forms, QA flows across pages. If you tried Playwright plus a raw LLM last year and watched your bill climb, this is worth a benchmark against your own tasks — because "fastest and cheapest" only means something on your workload, not theirs.
Eleven thousand stars in a week tells you people are hungry for agents that are actually deployable, not just demo-able.
If you want more repos like this each week, the newsletter link is in the description.
Hindsight: Agent Memory That Learns
Python · +10,944 stars this week
View on GitHub →What we said about hindsight
Here's one that caught my eye this week, and the name is honestly perfect. Hindsight tackles a problem most of us have hit the moment we build anything beyond a demo agent: memory that just accumulates instead of actually improving. Stuffing every past interaction into a vector store and calling it "memory" gets expensive and noisy fast. What Hindsight is going for is memory that learns — the agent reflects on what happened, keeps what's useful, and prunes the rest, so recall gets sharper over time instead of just bigger.
If you've wrestled with Mem0 or rolled your own retrieval layer on top of Pinecone, this is worth a look, especially since it's coming out of the Vectorize team who live in this space. It's Python, so it should slot into most existing stacks without much friction. Fair warning: eleven thousand stars in a week means it's early, so expect rough edges and shifting APIs.
If you want the full rundown on repos like this every week, the newsletter link is down in the description. Go grab it.
Google's open agentic orchestration runtime
Go · +10,085 stars this week
View on GitHub →What we said about ax
So Google just dropped "ax," and the numbers tell you people were waiting for this — ten thousand stars in a single week. At its core, ax is an orchestration runtime for agents, written in Go, which is already an interesting signal. Most of the agent frameworks we've covered live in Python land — LangGraph, CrewAI, that whole crowd. Google going with Go here suggests they're treating agent orchestration as infrastructure, not experimentation. Think long-running services, concurrency baked in, single-binary deploys, predictable memory. That's a different audience than someone prototyping in a notebook.
Where this matters is production. If you've ever tried to take a Python agent graph and make it survive real traffic, you know the pain. A Go runtime gives you the primitives to actually schedule, retry, and observe agents without duct tape. It's early, the ecosystem's thin, and Python still owns the tooling — so weigh that. But if you're a backend engineer eyeing agents seriously, ax is worth a real look.
We break down repos like this every week — grab the newsletter, link's in the description.
Shipping hard
A framework for building agentic apps
TypeScript · 4,216 commits this week
View on GitHub →What we said about agent-native
So Builder.io dropped "agent-native," and the framing here is genuinely different from most agent frameworks I've been poking at. Instead of treating the agent as a bolt-on layer that calls your app through an API, the idea is that the app itself is built around agentic behavior from the ground up — state, tools, and UI all designed to be driven by a model rather than just decorated with one. It's TypeScript-first, which matters if you're already living in a Next or Node world and don't want to context-switch into Python land just to wire up an agent loop.
Where this differs from something like LangChain or the various orchestration libraries is the emphasis on the *app* rather than the pipeline. Fewer abstractions between your intent and what ships to users. If you're a front-end or full-stack dev who's been curious about agents but bounced off the Python-heavy ecosystem, this is worth a real look.
Four thousand commits in a week tells me it's moving fast, so pin a version. If you want the weekly rundown of repos like this, the newsletter link's in the description.
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
TypeScript · 1,153 commits this week
View on GitHub →What we said about orca
Okay, let's talk about Orca from Stably AI, because this one's scratching an itch a lot of us have been feeling lately. If you've tried running more than one coding agent at once, you know the pain — juggling terminals, losing track of what each one's doing, no clean way to review their work. Orca calls itself an ADE, an agent development environment, and the pitch is basically an IDE-shaped cockpit for a whole fleet of agents running in parallel. What I appreciate is the "bring your own subscription" angle — you point it at agents you already pay for instead of getting locked into someone's token markup. And it runs on desktop, mobile, and a remote runtime, so you can kick off a task and check it from your phone. Compared to single-agent wrappers, this is aiming squarely at orchestration. Worth noting: eleven hundred commits this week, so it's moving fast — expect rough edges. If you like catching tools at this stage, the newsletter rounds up the best ones every week. Link's in the description.
Never stop coding. Free MIT AI gateway: one endpoint, 359 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by hundreds of contributors
TypeScript · 233 commits this week
View on GitHub →What we said about OmniRoute
OmniRoute is tackling a problem every one of us hits eventually — you're deep in a session, your provider throttles you, and suddenly you're stuck watching a rate-limit error instead of shipping. This is an MIT-licensed gateway that gives you a single endpoint fronting a huge spread of providers and models, with quota-aware fallback that quietly reroutes you when one runs dry. The part I find genuinely useful is the token compression layer — trimming context before it hits the model can meaningfully cut your bill on long agent runs, not just marginally.
Compared to something like OpenRouter, the pitch here is control: you self-host it, you own the routing logic, and it drops into Claude Code, Cursor, Cline, and Copilot without you rewiring your workflow. That makes it a solid fit for teams juggling free tiers or anyone tired of vendor lock-in.
Worth watching how the fallback holds up under real load, but the momentum is real.
If you want more finds like this before they blow up, the newsletter link's in the description.