Hacker News
Latest
Why the Bronze Age Collapsed
2026-09-30 @ 22:20:08Points: 8
The top secret URSALA, RAQUEL, and FARRAH satellites
2026-09-30 @ 22:03:04Points: 134Comments: 52
Gemini 4 Argon (High): Intelligence, Performance and Price Analysis
2026-09-30 @ 20:50:28Points: 84Comments: 45
Gitea 28.0
2026-09-30 @ 20:32:23Points: 65Comments: 31
Functional Ultrasound Imaging (fUSI) from scratch
2026-09-30 @ 20:06:39Points: 25Comments: 6
Gemini 4 Argon
2026-09-30 @ 20:04:37Points: 1001Comments: 672
CS240 AI Cheating Retrospective
2026-09-30 @ 19:54:29Points: 95Comments: 69
Dear Software Makers
2026-09-30 @ 19:45:38Points: 83Comments: 56
Halfspace experimental IDE for solid modeling with distance fields
2026-09-30 @ 19:44:37Points: 77Comments: 7
EDG C++ front-end goes public
2026-09-30 @ 19:26:37Points: 151Comments: 75
Surprisingly complex waves reveal the brain's inner workings
2026-09-30 @ 19:04:50Points: 121Comments: 40
Before pixels: Modular industrial dashboards
2026-09-30 @ 18:49:06Points: 51Comments: 11
5x faster Edge Functions: V8 isolates to Firecracker MicroVMs
2026-09-30 @ 18:17:45Points: 119Comments: 46
CHOMPI portable sampler instrument is now open-source (hardware and software)
2026-09-30 @ 17:42:13Points: 43Comments: 10
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
2026-09-30 @ 17:37:40Points: 126Comments: 57
We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case.
Inference engines today all make a performance tradeoff. They are either:
- Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4)
Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things.
Magnitude is built for maximum performance on your hardware and running local agents:
- On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels.
- Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance.
- Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run.
- Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concurrent sessions to share prefix caches, but optimize placement for memory-adjacency so single-session performance doesn't suffer.
Magnitude is fully open source (Apache 2.0). We built it in Rust, including a custom GPU kernel runtime and autotuner. We take inspiration from the best innovations in inference from academics (e.g. FlashAttention, FlashInfer, TurboQuant) as well as other engines (e.g. SGLang radix attention) to reach the performance ceiling.
Benchmarked against llama.cpp with Qwen 3.6 35B A3B (4 bit), 64k context, no speculative decoding:
Metal (Mac M4 Pro 48 GB) - 92% faster decode (30 tok/s → 57 tok/s) - 9% faster prefill (466 tok/s → 507 tok/s) - 28% less per-agent memory usage
CUDA (DGX Spark) - 19% faster decode (49 tok/s → 58 tok/s) - 23% faster prefill (2,033 tok/s → 2,507 tok/s) - 27% less per-agent memory usage
Magnitude ships as a desktop app that you can easily connect with whatever agents you already use (Pi, OpenCode, Hermes, Codex, and more). It automatically runs models on demand when these agents actually need them, and shuts them down after inactivity. Here's what it looks like: https://www.youtube.com/watch?v=0qE8BWEZu7o
We're excited to push Magnitude further to let you run bigger models on the same hardware while continuing to improve performance. Our plans include:
- Expert streaming: store experts on RAM or disk and load them just-in-time. This lets you run models bigger than what otherwise would fit on your GPU.
- Kernel compiler: our current kernels tune a few parameters to fit your hardware. We can take this further with a fully custom compiler that automatically chooses how to fuse kernels and which implementations to use, to make it fit to your hardware even better.
- Multi-device utilization: Make the best possible use of all hardware on a system (CPU, GPUs, RAM, disk) by detecting these and automatically solving for the best model layout.
We'd love for more people to try it out and give us feedback. Feel free to comment here, we'll be around all day!
Great Dirhombicosidodecahedron ("Miller's Monster")
2026-09-30 @ 17:00:44Points: 39Comments: 3
Bild AI (YC W25) Is Hiring a Founding Product Engineer
2026-09-30 @ 17:00:20Points: 1
Coltrane's Tone Circle
2026-09-30 @ 14:59:42Points: 28Comments: 7
Show HN: Lathoa, a math app for kids where the AI is wrong on purpose
2026-09-30 @ 14:38:57Points: 26Comments: 11
You can play one on the homepage without signing up.
The part that surprised me: it's hard to get an LLM to be wrong on purpose. Half the time it gives you the right answer and calls it wrong, or a "mistake" that's actually correct. So every case gets checked before a kid sees it. Where it can, a plain arithmetic check redoes the math exactly. A second model also solves the problem without seeing Errol's work. If anything disagrees, the case is thrown away.
The weak spot is that the second model can make the same mistake as the first. The arithmetic check is there for that, but it only works on English cases so far. German and Greek write decimals with a comma and I haven't got the parsing right yet.
What I'd really like to know: does finding someone else's mistake teach anything that solving the problem yourself doesn't? I'm not sure, and I'd like to hear from people who teach.