The best Hacker News stories from Show from the past day
Latest posts:
Show HN: Breadcrumb, record everything on your mac + context manager for AI
Hi HN, I'm Justin.<p>Breadcrumb records everything you do on your Mac (screen + meetings + AI transcripts + what you and your AI decided) and turns it into memory your AI can search. It's local and encrypted.<p>You can also teach it rules by talking to it and it makes sure the right rules turn up in the right context. Works with Claude Code / Codex / Cursor / opencode.<p>All of this is exposed to your AI as 30+ MCP tools (here's the definitions): <a href="https://innerloop.works/breadcrumb/mcp" rel="nofollow">https://innerloop.works/breadcrumb/mcp</a><p>I started it in June because I wanted to understand what was going on with my AI dev work. Then it just kept growing as I jammed on it 10 hours a day for 3 months.<p>I started by recording every screen so I could jump back to a moment and see what I'd been doing with the AI. Then I added MCP so the AI could search it. Then time tracking. Then meetings with transcripts and timestamped screenshots.<p>After that I started feeding the data back in. At first it was just a few relevant memories in the prompt. Then I started adding rules and rules modules.<p>Then I started recording the AI transcripts and all the other stuff the AI did (subagents, tools etc). I wanted a complete record so I could go back and see what the AI did and why.<p>I found the more data the AI had the more useful it became. One example. I got off a zoom call with the team at the day job and said:<p>"Review the meeting I just had, find all the bugs we discussed, and create JIRA tickets with screenshots."<p>It read the transcript. Found every bug we mentioned, pulled the screens from the screen timeline, spun up my local server, diagnosed each bug, reviewed the code, created 14 tickets in JIRA with diagnosis and likely fixes, attached 12 screenshots as evidence, and created a nice Slack message to send to my colleagues.<p>I never built a workflow for any of that. Claude had the meeting record, Breadcrumb tools, and rules I'd accumulated. It basically chained the steps together itself.<p>My favorite thing about the system is observability of my own work as an AI developer. I can trace a bug all the way back to the first thing I said (or the first thing the AI did) that caused it. Then I can add a rule so it doesn't happen again.<p>Anyway, my working system is simply Claude Desktop + Breadcrumb + Markdown. I find I don't need anything else.<p>It's local software for one person. Local models for transcription and diarization. Data encrypted with SQLCipher and a passphrase. You can exclude any apps, websites, words and phrases.<p>Breadcrumb itself only makes one network call which is a daily update check. (but of course your AI will use it and that is local or cloud depending on what you use).<p>It runs cool about 12% of one core on average (M series mac has 8+). You can set it to just save OCR text or to save screenshots. If just OCR it's about 12mb day (a lot of telemetry, indexing etc). If you record screenshots on meetings it's about 120mb/hour (1 every 2 secs). It can also record screenshots of everything that's heavy at about 160mb/day. It does local classification/naming with ollama with Gemma 3 4B (2.5 GB).<p>It's 100% free. There's no account you can just download and start using it. I'll monetize it via E2EE cloud sync and team sync. It's open beta. Just one developer working on it. Requires 16GB+ M Series Mac.<p>Thanks for reading!<p><a href="https://innerloop.works/breadcrumb" rel="nofollow">https://innerloop.works/breadcrumb</a> (homepage)<p><a href="https://innerloop.works/breadcrumb/download" rel="nofollow">https://innerloop.works/breadcrumb/download</a> (direct download)<p><a href="https://innerloop.works/breadcrumb/mcp" rel="nofollow">https://innerloop.works/breadcrumb/mcp</a> (mcp definition)
Show HN: Pyxel – A Python retro game engine with built-in art and sound editors
Hi HN, I'm the creator of Pyxel, a free, MIT-licensed retro game engine for Python. I've been developing it since 2018, and it recently passed 18,000 stars on GitHub.<p>Pyxel includes pixel art, tilemap, sound, and music editors, so you can create a game's visuals and audio as well as write its code in Python. The engine itself is implemented in Rust.<p>Games run on Windows, macOS, Linux, and the web, as well as supported Linux-based retro handhelds.<p>You can share a URL to launch a Pyxel game hosted on GitHub directly in a browser. You can also create games entirely in your browser with Pyxel Code Maker, or write chiptunes in Pyxel MML Studio.<p>See what people have made with Pyxel:
<a href="https://kitao.github.io/pyxel-user-examples/" rel="nofollow">https://kitao.github.io/pyxel-user-examples/</a><p>Try Code Maker:
<a href="https://kitao.github.io/pyxel/web/code-maker/" rel="nofollow">https://kitao.github.io/pyxel/web/code-maker/</a><p>I'm happy to answer questions about the engine and its implementation.
Show HN: Made an open-source Lego AI generator
Hi there :-) New on HN, first time posting.<p>Past year, around December, I started experimenting with making ChatGPT and Claude generate source code in LDraw language.<p>This LDraw is literally an "assembly" language, a low-level programming language that describes how to assemble LEGO pieces together into models, one placement instruction at a time.<p>When executed by specific tools, like e.g. LDView, LeoCAD, Studio... these instructions become LEGO CAD models, that can be interacted with, modified, etc.<p>Or, in other words: one LDraw source file in .mpd or .ldr format is equivalent to one LEGO CAD model.<p>So, the idea I had was: if I manage for maybe ChatGPT or Claude to generate high-quality LDraw source files... then, they would actually be generating high-quality LEGO CAD models, right?<p>Then, after months of iterations and trying one thing after the other... it worked!!!<p>Long story short: using GPT-6 Astra and Opus 5.5, I've managed to create a python toolset, instructions, and docs for agents in general. Now, these can be used by them to generate LDraw models.<p>I've packed it all as a dockerized web app for others to try and experiment, with several providers (and agents) to choose from: OpenAI, Claude and OpenRouter.<p>Here's a bunch of exmaples: <a href="https://anteloc.github.io/index-samples.html" rel="nofollow">https://anteloc.github.io/index-samples.html</a>.<p>If you are curious about the internals of an .mpd model, the "what was the agent picturing on its mind", open the .mpd file that got your attention on a text editor, and read the first line under the ones starting with "0 FILE".<p>I'd really appreciate feedback and comments, let's see where this goes =)
Show HN: Rhun, an open-source code editor written in assembly
I found that I'm not using even 1/3 of vim/vscode features anymore.<p>That's wht I'm building rhun - a small code editor for Linux, Windows and Apple silicon Macs. It obviously has Vim mode, a terminal, Git diffs and a panel for Claude Code or Codex sessions.<p>The editor and pixel renderer share an x86-64 assembly core. For Apple silicon, a build-time translator converts that core to AArch64, with separate platform adapters around it.
The latest release can draft commit messages using a local Ollama model or an existing Claude Code or Codex subscription.<p>It's a solo project, MIT licensed and still early.
Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia
Show HN: Audionaut – an open-source cross-platform multitrack audio editor
Show HN: Giving Opus 5.5 a simulated paint canvas
Show HN: Lathoa, a math app for kids where the AI is wrong on purpose
I made this for kids around 10 to 14. A robot called Errol solves a math problem step by step and one of the steps is wrong. The kid has to find it and say what's wrong with it. Sometimes nothing is wrong, so just saying "there's a mistake" every time doesn't work. The user needs to enter an explanation if she finds an error to gain more XP; speed matters also for more points. There is no direct interaction or chatting with an LLM. Lathoa's harness is stable and has many evaluation steps to catch inconsistencies and prompt injections.<p>You can play one on the homepage without signing up.<p>The part that surprised me: it's hard to get an LLM to be wrong on purpose. Half the time it gives you the right answer and calls it wrong, or a "mistake" that's actually correct. So every case gets checked before a kid sees it. Where it can, a plain arithmetic check redoes the math exactly. A second model also solves the problem without seeing Errol's work. If anything disagrees, the case is thrown away.<p>The weak spot is that the second model can make the same mistake as the first. The arithmetic check is there for that, but it only works on English cases so far. German and Greek write decimals with a comma and I haven't got the parsing right yet.<p>What I'd really like to know: does finding someone else's mistake teach anything that solving the problem yourself doesn't? I'm not sure, and I'd like to hear from people who teach.
Show HN: Open-source model routing for coding agents at Astra-level performance
A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.<p>First of all, a quick explanation: the Weave Router (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.<p>What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at <a href="https://weaveos.com/router">https://weaveos.com/router</a>!)<p>It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.<p>1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance.<p>Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.<p>In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but <i>how it got there</i>. We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach.<p>Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!<p>2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.<p>3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.<p>We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.<p>Our router is open source (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) so anyone can try it out. Or if you prefer you can use our hosted version (<a href="https://weaveos.com/router">https://weaveos.com/router</a>).
Show HN: Open-source model routing for coding agents at Astra-level performance
A few months ago we started building a model router for coding agents because we thought we could outperform any single model with an ensemble approach. Recently we’ve achieved that milestone and I want to talk about how we did it.<p>First of all, a quick explanation: the Weave Router (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) plugs into any coding agent (e.g. Claude Code or Codex) and intelligently switches between LLMs. So, for example, Astra handles tricky debugging or complex system design tasks, and Deepseek v4 Flash handles simple frontend updates.<p>What we’re announcing today is our new routing model, which we’re calling Weave Router 2.0. We benchmarked 2.0 against GPT-6 Astra on Terminal Bench 4.0 and SWE Atlas. On both benchmarks, the router had equivalent pass rates. On Terminal Bench, the router hit 52% of Astra’s cost, and completed tasks 2.2x faster. On SWE Atlas, the router cost 54% as much as Astra and ran 2.5x faster. (Full results on our website at <a href="https://weaveos.com/router">https://weaveos.com/router</a>!)<p>It turns out training a model to route effectively - taking into consideration model capabilities, costs, cache awareness, and more - is a really hard problem! I want to talk about three ways we were able to improve so much over the last few months: 1) a new architecture, 2) larger training data set size, and 3) smarter cache-eviction impact calculation.<p>1) a new architecture. Our initial approach used an RL model without many priors. While RL is still an important part of the story, the cost of fully exploring the space of routing decisions is very high, so we’ve taken some shortcuts that have significantly improved performance.<p>Consider how large the search space for the routing problem is. Take a typical coding agent session, with ~100 agent turns (i.e. 100 LLM API calls). Technically there are 100 chances to select a model. If we assume a roster of ~10 models (of course there are lots more but we can remove any that are Pareto dominated), then there are 10^100 possible paths through that session. We simply cannot explore all of them! So that's why clever tricks to shrink this space are so important.<p>In particular: we trained a hidden Markov model to trace the session state, then a classifier maps the session to one of a few buckets of similar models. Using the HMM allows us to evaluate not just where a session is currently, but <i>how it got there</i>. We've gotten significantly better performance on bucket selection by incorporating that information - we believe this is because two sessions that might look quite similar to a naive classifier are much better distinguished by this HMM approach.<p>Using this HMM + classifier to select a bucket first significantly shrinks the space to explore, by throwing out most models that could not reasonably serve the given session. This rearchitecture was the single biggest performance unlock!<p>2) larger training data set size (much less technically interesting but still an important part of the story). By using frontier LLMs to help us label a larger and more diverse set of coding agent sessions, we were able to bootstrap the two models discussed in 1) to a better state, while also providing even richer reward signals for RL.<p>3) smarter cache-eviction impact calculation. One of the hardest parts of routing well (if you care about saving money) is using the model caches intelligently. We built a subsystem that can calculate the expected value of switching models (and thus paying a high one-time cost to fill up a different cache) much more accurately, helping us avoid costly and unnecessary switches in more cases, while still switching when the benefit outweighs the cost. This is where most of our improvement on cost has come from.<p>We still have a lot of room to continue to improve (we won’t rest until we’re consistently beating Astra/Fable, not just tying!) but matching frontier model performance was a huge milestone for our routing model, and in my opinion validates our initial hypothesis that an ensemble of models can do better than any single model ever could.<p>Our router is open source (<a href="https://github.com/weave-os/router" rel="nofollow">https://github.com/weave-os/router</a>) so anyone can try it out. Or if you prefer you can use our hosted version (<a href="https://weaveos.com/router">https://weaveos.com/router</a>).
Vote on which of Hacker News' challenges for AI have been met
Show HN: TurboGPT: train 22KiB transformer in 13s
Inspired by minGPT for home experiments.<p>Requires CUDA 13.4; build script is Windows only.
Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
Show HN: Using 2D DFT, dithering, etc. to maximize eInk manga image quality
Kindle Comic Converter optimizes black & white (or color) comics and manga for E-ink ereaders like Kindle, Kobo, ReMarkable, and more. Pages display in fullscreen without margins, with proper fixed layout support. Output works great in KOReader.<p>Full details on what KCC does is in the readme with plenty of photos!<p>KCC has been under continuous development since 2012 and has code sign on both Windows and macOS. Also available on Linux.<p>-the current kcc dev
Show HN: Using 2D DFT, dithering, etc. to maximize eInk manga image quality
Kindle Comic Converter optimizes black & white (or color) comics and manga for E-ink ereaders like Kindle, Kobo, ReMarkable, and more. Pages display in fullscreen without margins, with proper fixed layout support. Output works great in KOReader.<p>Full details on what KCC does is in the readme with plenty of photos!<p>KCC has been under continuous development since 2012 and has code sign on both Windows and macOS. Also available on Linux.<p>-the current kcc dev
Show HN: Ledge.sh – Runnable Markdown Notes
Hi HN,<p>Ledge is a Markdown notebook that runs shell commands, code, SQL, etc from inside your own notes.<p>I built Ledge because I spend much of my day copy/pasting commands from my notes into the terminal. I was inspired by how much cmux helped me organize my terminals - but there was still a split brain between my notes and frequently run commands. I've been daily-driving it for the past few weeks and use it for deploys, API calls, smoke tests, etc.<p>Ledge runs your real shell just like a terminal app and can be hosted locally or remotely over SSH using ledge-server. I've been building it since July and have recently added support for all the major platforms: Mac (Silicon), Windows (WSL required), Linux, iOS, Android. Mobile devices require SSH access to a ledge-server and Android is still in beta and looking for beta testers (see link on website)!<p>It's built on Bun and Electrobun and is free and open-source. Feel free to review the code and contribute at <a href="https://github.com/ledgesh/ledge" rel="nofollow">https://github.com/ledgesh/ledge</a><p>Feedback is very much welcome. Any must-have features that are missing?
Show HN: Ledge.sh – Runnable Markdown Notes
Hi HN,<p>Ledge is a Markdown notebook that runs shell commands, code, SQL, etc from inside your own notes.<p>I built Ledge because I spend much of my day copy/pasting commands from my notes into the terminal. I was inspired by how much cmux helped me organize my terminals - but there was still a split brain between my notes and frequently run commands. I've been daily-driving it for the past few weeks and use it for deploys, API calls, smoke tests, etc.<p>Ledge runs your real shell just like a terminal app and can be hosted locally or remotely over SSH using ledge-server. I've been building it since July and have recently added support for all the major platforms: Mac (Silicon), Windows (WSL required), Linux, iOS, Android. Mobile devices require SSH access to a ledge-server and Android is still in beta and looking for beta testers (see link on website)!<p>It's built on Bun and Electrobun and is free and open-source. Feel free to review the code and contribute at <a href="https://github.com/ledgesh/ledge" rel="nofollow">https://github.com/ledgesh/ledge</a><p>Feedback is very much welcome. Any must-have features that are missing?
Show HN: Ledge.sh – Runnable Markdown Notes
Hi HN,<p>Ledge is a Markdown notebook that runs shell commands, code, SQL, etc from inside your own notes.<p>I built Ledge because I spend much of my day copy/pasting commands from my notes into the terminal. I was inspired by how much cmux helped me organize my terminals - but there was still a split brain between my notes and frequently run commands. I've been daily-driving it for the past few weeks and use it for deploys, API calls, smoke tests, etc.<p>Ledge runs your real shell just like a terminal app and can be hosted locally or remotely over SSH using ledge-server. I've been building it since July and have recently added support for all the major platforms: Mac (Silicon), Windows (WSL required), Linux, iOS, Android. Mobile devices require SSH access to a ledge-server and Android is still in beta and looking for beta testers (see link on website)!<p>It's built on Bun and Electrobun and is free and open-source. Feel free to review the code and contribute at <a href="https://github.com/ledgesh/ledge" rel="nofollow">https://github.com/ledgesh/ledge</a><p>Feedback is very much welcome. Any must-have features that are missing?
Show HN: A working 3D model of an Enigma machine
I watched the excellent Veritasium video [1] on the Enigma machine, and watched the full animation by Jared Owen [2], but was still a bit confused on how the inner mechanics of an Enigma machine work. I used Astra to build out the inner components through a combination of reference images, writing out hundreds of extremely detailed prompts, and building my own inspection tools to ensure that every part is sized and positioned in a historically accurate way.<p>It's still a work in progress, but would love any feedback on the experience so far!<p>1. <a href="https://www.youtube.com/watch?v=JsBZOcqZerk" rel="nofollow">https://www.youtube.com/watch?v=JsBZOcqZerk</a><p>2. <a href="https://www.youtube.com/watch?v=ybkkiGtJmkM" rel="nofollow">https://www.youtube.com/watch?v=ybkkiGtJmkM</a>
Show HN: A working 3D model of an Enigma machine
I watched the excellent Veritasium video [1] on the Enigma machine, and watched the full animation by Jared Owen [2], but was still a bit confused on how the inner mechanics of an Enigma machine work. I used Astra to build out the inner components through a combination of reference images, writing out hundreds of extremely detailed prompts, and building my own inspection tools to ensure that every part is sized and positioned in a historically accurate way.<p>It's still a work in progress, but would love any feedback on the experience so far!<p>1. <a href="https://www.youtube.com/watch?v=JsBZOcqZerk" rel="nofollow">https://www.youtube.com/watch?v=JsBZOcqZerk</a><p>2. <a href="https://www.youtube.com/watch?v=ybkkiGtJmkM" rel="nofollow">https://www.youtube.com/watch?v=ybkkiGtJmkM</a>