So I designed a local merge queue to have all commits land one at a time and fully tested. Hopefully this helps other folks with more modest machines. Appreciate any feedback.
Hacker News
Latest
NSF pilots 4-year PhDs with industry research placements
2026-07-30 @ 02:55:21Points: 65Comments: 66
Kuna: Decompiler Development in the Age of Coding Agents
2026-07-30 @ 02:41:16Points: 19Comments: 4
Logic for Programmers
2026-07-30 @ 00:51:39Points: 60Comments: 3
Show HN: A local merge queue for parallel Claude Code agents
2026-07-30 @ 00:20:16Points: 27Comments: 8
The Productivity Mirage
2026-07-29 @ 23:18:11Points: 156Comments: 49
Man and the Computer by John G. Kemeny (1972 book by the co-creator of BASIC)
2026-07-29 @ 22:53:44Points: 39Comments: 11
LLM Honeypot
2026-07-29 @ 22:51:03Points: 149Comments: 48
AI's top startups are barely publishing their research
2026-07-29 @ 21:25:40Points: 344Comments: 191
The Cold Email
2026-07-29 @ 21:06:42Points: 159Comments: 62
SalesPatriot (YC W25) Is Hiring FDEs
2026-07-29 @ 21:01:01Points: 1
The coolest use for the Vision Pro
2026-07-29 @ 20:39:40Points: 541Comments: 225
A Trampoline
2026-07-29 @ 20:14:36Points: 89Comments: 50
Kimi K3-256k
2026-07-29 @ 19:25:33Points: 398Comments: 116
Commodification of Intelligence: Good, Bad, and Ugly Circular AI Deals
2026-07-29 @ 18:57:10Points: 75Comments: 40
Turning a dumb AC unit smart (without losing my security deposit)
2026-07-29 @ 18:28:51Points: 142Comments: 113
Show HN: CheapFoodMap – A map of good meals under $10
2026-07-29 @ 16:59:52Points: 173Comments: 182
It's inspried by 거지맵 (Begger's Map) a Korean crowdsourced map students use to find cheap eats.
Ocverage is heaviest in Texas, since I live in Dallas, but have 1200 meals across 15 US cities. Seed data came from Google Review, 4.2 star or higher with at least 500 reviews, and verified price under $10 per menu item.
Things I would love feedback on : whether the price-freshness model makes sense, and what would make you trust the price on a site like this. How to encourage people to update prices, since inflation is making food price very frequent.
https://cheapfoodmap.com
Any and all suggestion will be super helpful. Thank you!
Some thoughts about Anthropic's new cryptanalysis results
2026-07-29 @ 16:42:20Points: 133Comments: 69
Keychron announces first open-source firmware for gaming mice
2026-07-29 @ 16:36:59Points: 342Comments: 140
Launch HN: Tokenless (YC S26) – Automatic model switching to save money
2026-07-29 @ 15:55:27Points: 58Comments: 50
The cost of AI tokens is top-of-mind for many. Companies like Uber and Salesforce have been complaining about blowing their yearly AI spend faster than expected.
Frontier models are amazing for dev work, but are so expensive. Open-source models are cheap and rapidly improving, closing the gap with frontier models, but aren’t quite there yet.
Tokenless gets you the best of both worlds–routing harder turns to smarter models only when needed, which keeps costs low.
Before Tokenless, I was doing a PhD at Princeton. While using coding/other agents, I constantly agonized over model choice, to make sure my AI spend was going as far as possible on my academic Cursor account.
At the same time, I was doing LLM research, and a small technique I developed while in recovery from NeurIPS submission season seemed to hit SOTA pretty fast. I was surprised that such simple ideas could do routing well.
We’ve been able to develop a version of the router that matches the performance of Claude Fable 5 at half the cost. The blog post on our website explores the technical details on how we did this (https://usetokenless.com/blog/building-tokenless/).
Highlights: - Our approach queries multiple models at once and uses their progress to make decisions (this technique is novel AFAIK, let us know if you know anyone else doing this). - Switching models doesn’t destroy the cache if the routing algorithm is aware of when the cache is hot/cold.
To come: - Adding Kimi K3, all other GPT efforts and more to the router
Go ahead and sign up on usetokenless.com and try using Tokenless with your agent, you’ll get $20 of free credit. Here’s a demo on how to use it: https://youtu.be/sjZWriclcls
Tokenless provides frontier-level intelligence for cheaper, so we’d love some feedback on how it feels to use, any corner cases that the router routes incorrectly, and whether you find the routing problem interesting!
Superlogical
2026-07-29 @ 15:41:33Points: 616Comments: 385
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
2026-07-29 @ 15:05:43Points: 726Comments: 252
I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal.
I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory.
The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are included.
The trick is to keep the shared part of the model and the KV cache in RAM, then stream only the routed experts needed for each token from SSD. An SSD is way slower than RAM, so the runtime uses a small expert cache and bounded parallel `pread`. While those reads are in flight, the GPU runs the shared part of the layer.
I ran more than 100 experiments. Most didn’t work. A few got me here. The experiments are described in the GitHub repo.
It currently generates 5–6 tok/s on an 8 GB M2 MacBook Air and 31–35 tok/s on an M5 MacBook Pro.
I also added an experimental OpenAI-compatible local server. It supports streaming and tool calls, and reuses one prompt prefix from the KV cache.
Try it! The Mac app is easy to install. On the first run, it will download 15 GB of weights from Hugging Face. The model is surprisingly capable.
I would love any kind of feedback!