The best Hacker News stories from Show from the past day
Latest posts:
Show HN: Lathoa, a math app for kids where the AI is wrong on purpose
I made this for kids around 10 to 14. A robot called Errol solves a math problem step by step and one of the steps is wrong. The kid has to find it and say what's wrong with it. Sometimes nothing is wrong, so just saying "there's a mistake" every time doesn't work. The user needs to enter an explanation if she finds an error to gain more XP; speed matters also for more points. There is no direct interaction or chatting with an LLM. Lathoa's harness is stable and has many evaluation steps to catch inconsistencies and prompt injections.<p>You can play one on the homepage without signing up.<p>The part that surprised me: it's hard to get an LLM to be wrong on purpose. Half the time it gives you the right answer and calls it wrong, or a "mistake" that's actually correct. So every case gets checked before a kid sees it. Where it can, a plain arithmetic check redoes the math exactly. A second model also solves the problem without seeing Errol's work. If anything disagrees, the case is thrown away.<p>The weak spot is that the second model can make the same mistake as the first. The arithmetic check is there for that, but it only works on English cases so far. German and Greek write decimals with a comma and I haven't got the parsing right yet.<p>What I'd really like to know: does finding someone else's mistake teach anything that solving the problem yourself doesn't? I'm not sure, and I'd like to hear from people who teach.
Show HN: Open-source model routing for coding agents at Astra-level performance
A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.<p>First of all, a quick explanation: the Weave Router (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.<p>What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at <a href="https://weaveos.com/router">https://weaveos.com/router</a>!)<p>It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.<p>1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance.<p>Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.<p>In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but <i>how it got there</i>. We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach.<p>Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!<p>2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.<p>3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.<p>We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.<p>Our router is open source (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) so anyone can try it out. Or if you prefer you can use our hosted version (<a href="https://weaveos.com/router">https://weaveos.com/router</a>).
Vote on which of Hacker News' challenges for AI have been met
Show HN: TurboGPT: train 22KiB transformer in 13s
Inspired by minGPT for home experiments.<p>Requires CUDA 13.4; build script is Windows only.
Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
Show HN: Using 2D DFT, dithering, etc. to maximize eInk manga image quality
Kindle Comic Converter optimizes black & white (or color) comics and manga for E-ink ereaders like Kindle, Kobo, ReMarkable, and more. Pages display in fullscreen without margins, with proper fixed layout support. Output works great in KOReader.<p>Full details on what KCC does is in the readme with plenty of photos!<p>KCC has been under continuous development since 2012 and has code sign on both Windows and macOS. Also available on Linux.<p>-the current kcc dev
Show HN: Using 2D DFT, dithering, etc. to maximize eInk manga image quality
Kindle Comic Converter optimizes black & white (or color) comics and manga for E-ink ereaders like Kindle, Kobo, ReMarkable, and more. Pages display in fullscreen without margins, with proper fixed layout support. Output works great in KOReader.<p>Full details on what KCC does is in the readme with plenty of photos!<p>KCC has been under continuous development since 2012 and has code sign on both Windows and macOS. Also available on Linux.<p>-the current kcc dev
Show HN: Ledge.sh – Runnable Markdown Notes
Hi HN,<p>Ledge is a Markdown notebook that runs shell commands, code, SQL, etc from inside your own notes.<p>I built Ledge because I spend much of my day copy/pasting commands from my notes into the terminal. I was inspired by how much cmux helped me organize my terminals - but there was still a split brain between my notes and frequently run commands. I've been daily-driving it for the past few weeks and use it for deploys, API calls, smoke tests, etc.<p>Ledge runs your real shell just like a terminal app and can be hosted locally or remotely over SSH using ledge-server. I've been building it since July and have recently added support for all the major platforms: Mac (Silicon), Windows (WSL required), Linux, iOS, Android. Mobile devices require SSH access to a ledge-server and Android is still in beta and looking for beta testers (see link on website)!<p>It's built on Bun and Electrobun and is free and open-source. Feel free to review the code and contribute at <a href="https://github.com/ledgesh/ledge" rel="nofollow">https://github.com/ledgesh/ledge</a><p>Feedback is very much welcome. Any must-have features that are missing?
Show HN: Ledge.sh – Runnable Markdown Notes
Hi HN,<p>Ledge is a Markdown notebook that runs shell commands, code, SQL, etc from inside your own notes.<p>I built Ledge because I spend much of my day copy/pasting commands from my notes into the terminal. I was inspired by how much cmux helped me organize my terminals - but there was still a split brain between my notes and frequently run commands. I've been daily-driving it for the past few weeks and use it for deploys, API calls, smoke tests, etc.<p>Ledge runs your real shell just like a terminal app and can be hosted locally or remotely over SSH using ledge-server. I've been building it since July and have recently added support for all the major platforms: Mac (Silicon), Windows (WSL required), Linux, iOS, Android. Mobile devices require SSH access to a ledge-server and Android is still in beta and looking for beta testers (see link on website)!<p>It's built on Bun and Electrobun and is free and open-source. Feel free to review the code and contribute at <a href="https://github.com/ledgesh/ledge" rel="nofollow">https://github.com/ledgesh/ledge</a><p>Feedback is very much welcome. Any must-have features that are missing?
Show HN: A working 3D model of an Enigma machine
I watched the excellent Veritasium video [1] on the Enigma machine, and watched the full animation by Jared Owen [2], but was still a bit confused on how the inner mechanics of an Enigma machine work. I used Astra to build out the inner components through a combination of reference images, writing out hundreds of extremely detailed prompts, and building my own inspection tools to ensure that every part is sized and positioned in a historically accurate way.<p>It's still a work in progress, but would love any feedback on the experience so far!<p>1. <a href="https://www.youtube.com/watch?v=JsBZOcqZerk" rel="nofollow">https://www.youtube.com/watch?v=JsBZOcqZerk</a><p>2. <a href="https://www.youtube.com/watch?v=ybkkiGtJmkM" rel="nofollow">https://www.youtube.com/watch?v=ybkkiGtJmkM</a>
Show HN: A working 3D model of an Enigma machine
I watched the excellent Veritasium video [1] on the Enigma machine, and watched the full animation by Jared Owen [2], but was still a bit confused on how the inner mechanics of an Enigma machine work. I used Astra to build out the inner components through a combination of reference images, writing out hundreds of extremely detailed prompts, and building my own inspection tools to ensure that every part is sized and positioned in a historically accurate way.<p>It's still a work in progress, but would love any feedback on the experience so far!<p>1. <a href="https://www.youtube.com/watch?v=JsBZOcqZerk" rel="nofollow">https://www.youtube.com/watch?v=JsBZOcqZerk</a><p>2. <a href="https://www.youtube.com/watch?v=ybkkiGtJmkM" rel="nofollow">https://www.youtube.com/watch?v=ybkkiGtJmkM</a>
Show HN: JBR-001 – An open-source 3D printable desktop robot
We built a desktop companion robot that is powered by Arduino UNO Q. Project is open source: source code, 3D printable files and assembly instructions are available at Arduino Project Hub.<p>Robot is equipped with camera, distance sensor, buzzer, moves head and arms with 3 servo motors and has an animated display. It can recognise objects and respond to it. We build it with idea this should be fun project people can 3D print and build at home themself.
Show HN: JBR-001 – An open-source 3D printable desktop robot
We built a desktop companion robot that is powered by Arduino UNO Q. Project is open source: source code, 3D printable files and assembly instructions are available at Arduino Project Hub.<p>Robot is equipped with camera, distance sensor, buzzer, moves head and arms with 3 servo motors and has an animated display. It can recognise objects and respond to it. We build it with idea this should be fun project people can 3D print and build at home themself.
Show HN: Dental Scope – Interactive 3D dental anatomy
Show HN: Dental Scope – Interactive 3D dental anatomy
Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
Benchmarks how well Harness+models can create a Pac-Man game from a single prompt:<p>“Create a Pac-Man game in a single HTML page”<p>Each model gets one shot — no follow-up prompts or fixes.
Show HN: Real-time Solar System with 526k asteroids and all tracked satellites
Show HN: Real-time Solar System with 526k asteroids and all tracked satellites
Show HN: Real-time Solar System with 526k asteroids and all tracked satellites
Show HN: NSL – WSL for Linux
One of the things that Windows really got right is WSL2. I drive an atomic Linux distro for daily use, but wanted a way to develop with multiple different distros with that same WSL UX. NSL is my answer. It is a faithful reproduction of the developer experience, powered by a single VM that hosts one or more systemd-nspawn containers with your development instances. Host file edits and port sharing come along for the ride, just like WSL. Take a look and tell me what you think... It's yet another step in my long journey to keep my host installation free from all the changing and breaking dev dependencies that force a reinstall every few months.