The best Hacker News stories from Show from the past day

Go back

Latest posts:

Show HN: Our space game has a built-in RISC-V emulator that runs Linux

(Edit: original URL was <a href="https://againstallodds.games/" rel="nofollow">https://againstallodds.games/</a>, but we've switched it to <a href="https://againstallodds.games/blog/2026/10/03/our-risc-v-emulator-pasriscv/" rel="nofollow">https://againstallodds.games/blog/2026/10/03/our-risc-v-emul...</a> in response to user requests for explanation.)<p>We're a tiny indie studio, all with a background in the demoscene and for about 3 years now we're developing a space planet terraforming game called SEEDS - Echoes Beneath the Sands.<p>Our protagonist Naxiah is stranded on a small, desolated planet. You're working for an intergalactic distributor of seeds and terraforming equipment, to kickstart new planets in far away galaxies. You must deliver seeds and utilities to an unchartered region, but on your way, you crash on a small planet. Luckily, you have some equipment with you in the ship, so you're able to spawn a base and survive. But for how long? And is the planet really deserted..? :) Well, of course not, there are aliens and other space creatures. And of course, there's ROBO RB-23, which accompanies your adventure.<p>The game is written in Object Pascal and uses our own game engine PasVulkan, our own physics engine, as well as our own RISC-V 64-bit emulator PasRISCV. We love to build stuff, a lot of these things are FOSS.<p>The emulator is quite complete (even with RVV, but that's way too slow emulated to be useful) and runs a stock kernel (6.18.3 as of today). We're using Alpine Linux for our base distribution. All our in-game programs to interact with the planet are native RISC-V Linux programs.<p>We also have our own (free and open source) scripting language POCA (class and prototyped based, JS/Lua inspired, which makes it quite easy for us to script things like HUD, animal behaviour, our story flow graph, etc.<p>We have most of the wiring ready for players to be able to program and reshape and populate the whole planet in automation, although it's not our primary concept for the game (building stuff by hand, making the planet beautiful is also fun). As it's a normal Linux system, you'll be able to use Ruby, Python, or even Rust, C, C++, and Go to change and automate your world.<p>We're also thinking of making a bunch of small games which you can play inside the game, in the base on a computer, on an arcade machine, or in the spaceship. And of course, it also runs DOOM. We already prototyped a small 2D space shooter and a small arcade game, both playable within the game. We also feature a table with hologram games to be played, e.g. chess and a certain block game. We also plan to make this hologram game engine available from the computer, which means, you'll even be able to program own hologram style games.<p>We generally want to make the in-game computer more hackable in future, so that people can create and share their own little games.<p>Currently, it's only a sandbox game (with beginnings of automation), to build stuff with our pre-made items. You can also assemble these to other, ready-made and shareable bigger items. This is the foundation for us to further develop survival and complete our story mode.<p>Our game is still under active development, and we finally got our Steam page ready. We have tons of other things planned, but right now, it's more about polishing and quality for the Early Access.<p>You can watch some videos and screenshots on our website. And we'd love to read your thoughts about his. It has been quite a ride, so far..

Show HN: Pi pod – Run your pi coding agent in sandboxes on your own server

pi pod runs sessions of the pi coding agent in isolated sandboxes ("pods") on a server you run, in composable environments.<p>----<p>Since moving my company towards AI-native work, I have been really frustrated by the state of "agentic engineering" environments. Products by the labs (claude code, codex) lock you into a single provider for your tokens. Agnostic solutions (factory, devin, arguably cursor) make you pay per-token costs. None of these products allow you to fully customize the harness, and of course they all run on someone else's infrastructure.<p>I've been an early and fervent user of pi, which I think is fantastically simple and beautiful software. I have felt it needs an environment for it to work across platforms with fully functional composability for teams.<p>This is very much a work in progress, but for my team this has been a much needed solution and has helped us tremendously. I hope you will give it a try and let me know how you would like it to improve.

Show HN: Offrun – manage every coding agent from one workspace

Run Claude Code, Codex, AGY, and Grok Build side by side. See who is working, who needs you, and what every account has left.

Aleph Alpha Kolibri: How the sovereign German LLM works

<a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/" rel="nofollow">https://aleph-alpha.com/en/blog/kolibri-has-landed-a-soverei...</a>

Show HN: Breadcrumb, record everything on your mac + context manager for AI

Hi HN, I'm Justin.<p>Breadcrumb records everything you do on your Mac (screen + meetings + AI transcripts + what you and your AI decided) and turns it into memory your AI can search. It's local and encrypted.<p>You can also teach it rules by talking to it and it makes sure the right rules turn up in the right context. Works with Claude Code / Codex / Cursor / opencode.<p>All of this is exposed to your AI as 30+ MCP tools (here's the definitions): <a href="https://innerloop.works/breadcrumb/mcp" rel="nofollow">https://innerloop.works/breadcrumb/mcp</a><p>I started it in June because I wanted to understand what was going on with my AI dev work. Then it just kept growing as I jammed on it 10 hours a day for 3 months.<p>I started by recording every screen so I could jump back to a moment and see what I'd been doing with the AI. Then I added MCP so the AI could search it. Then time tracking. Then meetings with transcripts and timestamped screenshots.<p>After that I started feeding the data back in. At first it was just a few relevant memories in the prompt. Then I started adding rules and rules modules.<p>Then I started recording the AI transcripts and all the other stuff the AI did (subagents, tools etc). I wanted a complete record so I could go back and see what the AI did and why.<p>I found the more data the AI had the more useful it became. One example. I got off a zoom call with the team at the day job and said:<p>"Review the meeting I just had, find all the bugs we discussed, and create JIRA tickets with screenshots."<p>It read the transcript. Found every bug we mentioned, pulled the screens from the screen timeline, spun up my local server, diagnosed each bug, reviewed the code, created 14 tickets in JIRA with diagnosis and likely fixes, attached 12 screenshots as evidence, and created a nice Slack message to send to my colleagues.<p>I never built a workflow for any of that. Claude had the meeting record, Breadcrumb tools, and rules I'd accumulated. It basically chained the steps together itself.<p>My favorite thing about the system is observability of my own work as an AI developer. I can trace a bug all the way back to the first thing I said (or the first thing the AI did) that caused it. Then I can add a rule so it doesn't happen again.<p>Anyway, my working system is simply Claude Desktop + Breadcrumb + Markdown. I find I don't need anything else.<p>It's local software for one person. Local models for transcription and diarization. Data encrypted with SQLCipher and a passphrase. You can exclude any apps, websites, words and phrases.<p>Breadcrumb itself only makes one network call which is a daily update check. (but of course your AI will use it and that is local or cloud depending on what you use).<p>It runs cool about 12% of one core on average (M series mac has 8+). You can set it to just save OCR text or to save screenshots. If just OCR it's about 12mb day (a lot of telemetry, indexing etc). If you record screenshots on meetings it's about 120mb/hour (1 every 2 secs). It can also record screenshots of everything that's heavy at about 160mb/day. It does local classification/naming with ollama with Gemma 3 4B (2.5 GB).<p>It's 100% free. There's no account you can just download and start using it. I'll monetize it via E2EE cloud sync and team sync. It's open beta. Just one developer working on it. Requires 16GB+ M Series Mac.<p>Thanks for reading!<p><a href="https://innerloop.works/breadcrumb" rel="nofollow">https://innerloop.works/breadcrumb</a> (homepage)<p><a href="https://innerloop.works/breadcrumb/download" rel="nofollow">https://innerloop.works/breadcrumb/download</a> (direct download)<p><a href="https://innerloop.works/breadcrumb/mcp" rel="nofollow">https://innerloop.works/breadcrumb/mcp</a> (mcp definition)

Show HN: Pyxel – A Python retro game engine with built-in art and sound editors

Hi HN, I'm the creator of Pyxel, a free, MIT-licensed retro game engine for Python. I've been developing it since 2018, and it recently passed 18,000 stars on GitHub.<p>Pyxel includes pixel art, tilemap, sound, and music editors, so you can create a game's visuals and audio as well as write its code in Python. The engine itself is implemented in Rust.<p>Games run on Windows, macOS, Linux, and the web, as well as supported Linux-based retro handhelds.<p>You can share a URL to launch a Pyxel game hosted on GitHub directly in a browser. You can also create games entirely in your browser with Pyxel Code Maker, or write chiptunes in Pyxel MML Studio.<p>See what people have made with Pyxel: <a href="https://kitao.github.io/pyxel-user-examples/" rel="nofollow">https://kitao.github.io/pyxel-user-examples/</a><p>Try Code Maker: <a href="https://kitao.github.io/pyxel/web/code-maker/" rel="nofollow">https://kitao.github.io/pyxel/web/code-maker/</a><p>I'm happy to answer questions about the engine and its implementation.

Show HN: Pyxel – A Python retro game engine with built-in art and sound editors

Hi HN, I'm the creator of Pyxel, a free, MIT-licensed retro game engine for Python. I've been developing it since 2018, and it recently passed 18,000 stars on GitHub.<p>Pyxel includes pixel art, tilemap, sound, and music editors, so you can create a game's visuals and audio as well as write its code in Python. The engine itself is implemented in Rust.<p>Games run on Windows, macOS, Linux, and the web, as well as supported Linux-based retro handhelds.<p>You can share a URL to launch a Pyxel game hosted on GitHub directly in a browser. You can also create games entirely in your browser with Pyxel Code Maker, or write chiptunes in Pyxel MML Studio.<p>See what people have made with Pyxel: <a href="https://kitao.github.io/pyxel-user-examples/" rel="nofollow">https://kitao.github.io/pyxel-user-examples/</a><p>Try Code Maker: <a href="https://kitao.github.io/pyxel/web/code-maker/" rel="nofollow">https://kitao.github.io/pyxel/web/code-maker/</a><p>I'm happy to answer questions about the engine and its implementation.

Show HN: Made an open-source Lego AI generator

Hi there :-) New on HN, first time posting.<p>Past year, around December, I started experimenting with making ChatGPT and Claude generate source code in LDraw language.<p>This LDraw is literally an "assembly" language, a low-level programming language that describes how to assemble LEGO pieces together into models, one placement instruction at a time.<p>When executed by specific tools, like e.g. LDView, LeoCAD, Studio... these instructions become LEGO CAD models, that can be interacted with, modified, etc.<p>Or, in other words: one LDraw source file in .mpd or .ldr format is equivalent to one LEGO CAD model.<p>So, the idea I had was: if I manage for maybe ChatGPT or Claude to generate high-quality LDraw source files... then, they would actually be generating high-quality LEGO CAD models, right?<p>Then, after months of iterations and trying one thing after the other... it worked!!!<p>Long story short: using GPT-6 Astra and Opus 5.5, I've managed to create a python toolset, instructions, and docs for agents in general. Now, these can be used by them to generate LDraw models.<p>I've packed it all as a dockerized web app for others to try and experiment, with several providers (and agents) to choose from: OpenAI, Claude and OpenRouter.<p>Here's a bunch of exmaples: <a href="https://anteloc.github.io/index-samples.html" rel="nofollow">https://anteloc.github.io/index-samples.html</a>.<p>If you are curious about the internals of an .mpd model, the "what was the agent picturing on its mind", open the .mpd file that got your attention on a text editor, and read the first line under the ones starting with "0 FILE".<p>I'd really appreciate feedback and comments, let's see where this goes =)

Show HN: Made an open-source Lego AI generator

Hi there :-) New on HN, first time posting.<p>Past year, around December, I started experimenting with making ChatGPT and Claude generate source code in LDraw language.<p>This LDraw is literally an "assembly" language, a low-level programming language that describes how to assemble LEGO pieces together into models, one placement instruction at a time.<p>When executed by specific tools, like e.g. LDView, LeoCAD, Studio... these instructions become LEGO CAD models, that can be interacted with, modified, etc.<p>Or, in other words: one LDraw source file in .mpd or .ldr format is equivalent to one LEGO CAD model.<p>So, the idea I had was: if I manage for maybe ChatGPT or Claude to generate high-quality LDraw source files... then, they would actually be generating high-quality LEGO CAD models, right?<p>Then, after months of iterations and trying one thing after the other... it worked!!!<p>Long story short: using GPT-6 Astra and Opus 5.5, I've managed to create a python toolset, instructions, and docs for agents in general. Now, these can be used by them to generate LDraw models.<p>I've packed it all as a dockerized web app for others to try and experiment, with several providers (and agents) to choose from: OpenAI, Claude and OpenRouter.<p>Here's a bunch of exmaples: <a href="https://anteloc.github.io/index-samples.html" rel="nofollow">https://anteloc.github.io/index-samples.html</a>.<p>If you are curious about the internals of an .mpd model, the "what was the agent picturing on its mind", open the .mpd file that got your attention on a text editor, and read the first line under the ones starting with "0 FILE".<p>I'd really appreciate feedback and comments, let's see where this goes =)

Show HN: Rhun, an open-source code editor written in assembly

I found that I'm not using even 1/3 of vim/vscode features anymore.<p>That's wht I'm building rhun - a small code editor for Linux, Windows and Apple silicon Macs. It obviously has Vim mode, a terminal, Git diffs and a panel for Claude Code or Codex sessions.<p>The editor and pixel renderer share an x86-64 assembly core. For Apple silicon, a build-time translator converts that core to AArch64, with separate platform adapters around it. The latest release can draft commit messages using a local Ollama model or an existing Claude Code or Codex subscription.<p>It's a solo project, MIT licensed and still early.

Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

Show HN: Audionaut – an open-source cross-platform multitrack audio editor

Show HN: Audionaut – an open-source cross-platform multitrack audio editor

Show HN: Giving Opus 5.5 a simulated paint canvas

Show HN: Giving Opus 5.5 a simulated paint canvas

Show HN: Lathoa, a math app for kids where the AI is wrong on purpose

I made this for kids around 10 to 14. A robot called Errol solves a math problem step by step and one of the steps is wrong. The kid has to find it and say what's wrong with it. Sometimes nothing is wrong, so just saying "there's a mistake" every time doesn't work. The user needs to enter an explanation if she finds an error to gain more XP; speed matters also for more points. There is no direct interaction or chatting with an LLM. Lathoa's harness is stable and has many evaluation steps to catch inconsistencies and prompt injections.<p>You can play one on the homepage without signing up.<p>The part that surprised me: it's hard to get an LLM to be wrong on purpose. Half the time it gives you the right answer and calls it wrong, or a "mistake" that's actually correct. So every case gets checked before a kid sees it. Where it can, a plain arithmetic check redoes the math exactly. A second model also solves the problem without seeing Errol's work. If anything disagrees, the case is thrown away.<p>The weak spot is that the second model can make the same mistake as the first. The arithmetic check is there for that, but it only works on English cases so far. German and Greek write decimals with a comma and I haven't got the parsing right yet.<p>What I'd really like to know: does finding someone else's mistake teach anything that solving the problem yourself doesn't? I'm not sure, and I'd like to hear from people who teach.

Show HN: Open-source model routing for coding agents at Astra-level performance

A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.<p>First of all, a quick explanation: the Weave Router (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.<p>What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at <a href="https://weaveos.com/router">https://weaveos.com/router</a>!)<p>It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.<p>1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance.<p>Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.<p>In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but <i>how it got there</i>. We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach.<p>Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!<p>2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.<p>3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.<p>We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.<p>Our router is open source (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) so anyone can try it out. Or if you prefer you can use our hosted version (<a href="https://weaveos.com/router">https://weaveos.com/router</a>).

Show HN: Open-source model routing for coding agents at Astra-level performance

A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.<p>First of all, a quick explanation: the Weave Router (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.<p>What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at <a href="https://weaveos.com/router">https://weaveos.com/router</a>!)<p>It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.<p>1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance.<p>Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.<p>In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but <i>how it got there</i>. We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach.<p>Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!<p>2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.<p>3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.<p>We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.<p>Our router is open source (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) so anyone can try it out. Or if you prefer you can use our hosted version (<a href="https://weaveos.com/router">https://weaveos.com/router</a>).

Show HN: Open-source model routing for coding agents at Astra-level performance

A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.<p>First of all, a quick explanation: the Weave Router (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.<p>What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at <a href="https://weaveos.com/router">https://weaveos.com/router</a>!)<p>It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.<p>1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance.<p>Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.<p>In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but <i>how it got there</i>. We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach.<p>Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!<p>2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.<p>3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.<p>We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.<p>Our router is open source (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) so anyone can try it out. Or if you prefer you can use our hosted version (<a href="https://weaveos.com/router">https://weaveos.com/router</a>).

1 2 3 ... 1051 1052 1053 >