Hacker News

Latest

Show HN: Yantra – an LALR(1) parser generator for C++

2026-10-01 @ 02:30:48Points: 30Comments: 15

Most LALR parser generators (Yacc, Bison, Lemon) run your semantic actions during parsing, as each rule reduces, bottom-up.

That means at the time a rule's action runs, you don't yet know what its parent looks like. This pushes a lot of grammars toward hand-built AST classes and a separate walking pass whenever you need to look ahead into siblings or defer a decision until more context is available.

On the other hand, Yantra always builds the whole AST first, then walks it top-down in a separate pass, calling your semantic actions as it goes. A parent rule's action can run before its children are visited.

A single grammar can define more than one walker. For example, one that emits C++, another that emits Java, from the same parse. The AST and the walker classes are both generated for you.

A small example (full version, with compile commands, in the README):

  start := expr;

  expr := expr(a) PLUS expr(b)
  %{
      std::cout << "Adding" << std::endl;
  %}

  expr := NUMBER(N)
  %{
      std::cout << "Number: " << N.text << std::endl;
  %}

  NUMBER := "\d+";
  PLUS := "\+";
  WS := "\s+"!;
Running this on "1 + 2 + 3" prints:

  Adding
  Number: 1
  Adding
  Number: 2
  Number: 3
The outer "Adding", the root of the tree, prints first, before either of its children. That's only possible because the whole tree exists before any action runs.

Some other things about it: integrated lexer with mode support (for things like nested comments), an optional amalgamated single-file output mode with a generated main(), C++23, MIT licensed.

It's young (0.5.1, pre-1.0) and single-maintainer, so treat it as early. I'd rather know what breaks than have it look more finished than it is.

Known gaps are listed at https://github.com/TantrixAuto/yantra/blob/main/docs/known_l...

Repo: https://github.com/TantrixAuto/yantra

Feedback and questions are all welcome. I'll be around.

56k.rip – the 1996 dial-up internet experience

2026-09-30 @ 22:08:50Points: 240Comments: 98

The top secret URSALA, RAQUEL, and FARRAH satellites (2025)

2026-09-30 @ 22:03:04Points: 273Comments: 134

Functional Ultrasound Imaging (fUSI) from scratch

2026-09-30 @ 20:06:39Points: 50Comments: 10

Gemini 4 Argon

2026-09-30 @ 20:04:37Points: 1553Comments: 1029

Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236

Halfspace experimental IDE for solid modeling with distance fields

2026-09-30 @ 19:44:37Points: 170Comments: 10

EDG C++ front-end goes public

2026-09-30 @ 19:26:37Points: 233Comments: 122

Surprisingly complex waves reveal the brain's inner workings

2026-09-30 @ 19:04:50Points: 228Comments: 91

Before pixels: Modular industrial dashboards

2026-09-30 @ 18:49:06Points: 244Comments: 44

5x faster Edge Functions: V8 isolates to Firecracker MicroVMs

2026-09-30 @ 18:17:45Points: 205Comments: 89

CHOMPI portable sampler instrument is now open-source (hardware and software)

2026-09-30 @ 17:42:13Points: 99Comments: 27

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

2026-09-30 @ 17:37:40Points: 180Comments: 88

We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case.

Inference engines today all make a performance tradeoff. They are either:

- Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4)

Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things.

Magnitude is built for maximum performance on your hardware and running local agents:

- On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels.

- Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance.

- Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run.

- Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concurrent sessions to share prefix caches, but optimize placement for memory-adjacency so single-session performance doesn't suffer.

Magnitude is fully open source (Apache 2.0). We built it in Rust, including a custom GPU kernel runtime and autotuner. We take inspiration from the best innovations in inference from academics (e.g. FlashAttention, FlashInfer, TurboQuant) as well as other engines (e.g. SGLang radix attention) to reach the performance ceiling.

Benchmarked against llama.cpp with Qwen 3.6 35B A3B (4 bit), 64k context, no speculative decoding:

Metal (Mac M4 Pro 48 GB) - 92% faster decode (30 tok/s → 57 tok/s) - 9% faster prefill (466 tok/s → 507 tok/s) - 28% less per-agent memory usage

CUDA (DGX Spark) - 19% faster decode (49 tok/s → 58 tok/s) - 23% faster prefill (2,033 tok/s → 2,507 tok/s) - 27% less per-agent memory usage

Magnitude ships as a desktop app that you can easily connect with whatever agents you already use (Pi, OpenCode, Hermes, Codex, and more). It automatically runs models on demand when these agents actually need them, and shuts them down after inactivity. Here's what it looks like: https://www.youtube.com/watch?v=0qE8BWEZu7o

We're excited to push Magnitude further to let you run bigger models on the same hardware while continuing to improve performance. Our plans include:

- Expert streaming: store experts on RAM or disk and load them just-in-time. This lets you run models bigger than what otherwise would fit on your GPU.

- Kernel compiler: our current kernels tune a few parameters to fit your hardware. We can take this further with a fully custom compiler that automatically chooses how to fuse kernels and which implementations to use, to make it fit to your hardware even better.

- Multi-device utilization: Make the best possible use of all hardware on a system (CPU, GPUs, RAM, disk) by detecting these and automatically solving for the best model layout.

We'd love for more people to try it out and give us feedback. Feel free to comment here, we'll be around all day!

Great Dirhombicosidodecahedron ("Miller's Monster")

2026-09-30 @ 17:00:44Points: 77Comments: 17

Coltrane's Tone Circle

2026-09-30 @ 14:59:42Points: 77Comments: 22

SDF Public Access Unix System ... est. 1987

2026-09-30 @ 14:36:02Points: 97Comments: 25

A brief history of the Bloomberg terminal

2026-09-30 @ 14:34:07Points: 344Comments: 144

What TLA+ can and can't check

2026-09-30 @ 13:57:06Points: 223Comments: 47

SDF vs. MSDF vs. Slug: GPU Text Rendering

2026-09-30 @ 13:50:50Points: 168Comments: 62

The last time my family was replaced by technology

2026-09-30 @ 13:06:11Points: 298Comments: 582

You said no MCP

2026-09-30 @ 09:55:23Points: 652Comments: 358

Singapore govt dating app uses Gale-Shapley stable marriage algorithm

2026-09-30 @ 09:27:04Points: 432Comments: 425

OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network

2026-09-30 @ 08:43:21Points: 195Comments: 95

Doing a Machine Learning PhD While Working in Japan

2026-09-30 @ 07:33:45Points: 126Comments: 41

LinkedIn Larpmaxxing

2026-09-30 @ 04:19:57Points: 271Comments: 223

Responsible Release of AI-Generated Mathematics

2026-09-30 @ 02:36:12Points: 109Comments: 155

Show HN: Ledge.sh – Runnable Markdown Notes

2026-09-29 @ 23:41:34Points: 185Comments: 82

Ledge is a Markdown notebook that runs shell commands, code, SQL, etc from inside your own notes.

I built Ledge because I spend much of my day copy/pasting commands from my notes into the terminal. I was inspired by how much cmux helped me organize my terminals - but there was still a split brain between my notes and frequently run commands. I've been daily-driving it for the past few weeks and use it for deploys, API calls, smoke tests, etc.

Ledge runs your real shell just like a terminal app and can be hosted locally or remotely over SSH using ledge-server. I've been building it since July and have recently added support for all the major platforms: Mac (Silicon), Windows (WSL required), Linux, iOS, Android. Mobile devices require SSH access to a ledge-server and Android is still in beta and looking for beta testers (see link on website)!

It's built on Bun and Electrobun and is free and open-source. Feel free to review the code and contribute at https://github.com/ledgesh/ledge

Feedback is very much welcome. Any must-have features that are missing?

PlayBook: A Programmable Paper Notebook [video]

2026-09-29 @ 15:40:03Points: 31Comments: 6

Book of Shapes – Collection of minimal, generative and customizable SVG-patterns

2026-09-29 @ 14:56:22Points: 184Comments: 13

Why the Bronze Age Collapsed

2026-09-29 @ 10:10:18Points: 338Comments: 223

Burning Man death rates – A short lesson in statistics

2026-09-28 @ 18:57:16Points: 163Comments: 214

Archives

2026

2025

2024

2023

2022