Hacker News

Latest

Alphabet's cash burn raises alarm for Big Tech as AI spending climbs

2026-07-23 @ 13:10:02Points: 162Comments: 152

AI Companies Are Trying to Hide a Staggering Amount of Debt

2026-07-23 @ 13:09:10Points: 132Comments: 63

OpenAI and Anthropic unite against open-weight AI risks to their bottom line

2026-07-23 @ 13:00:03Points: 213Comments: 245

EU fines Google €890M for competition breaches over search and apps

2026-07-23 @ 10:07:22Points: 139Comments: 146

Cruller: Bun's Zig Runtime, Continued on Zig 0.16

2026-07-23 @ 05:40:20Points: 121Comments: 79

Amiga 1000: Ten years ahead of its time

2026-07-23 @ 05:24:17Points: 135Comments: 120

Protecting our FLOSS commons from LLMs

2026-07-23 @ 01:14:53Points: 149Comments: 86

Medici family mystery may be solved after more than 400 years

2026-07-22 @ 21:55:34Points: 141Comments: 45

Malleable Computing, Emacs, and You

2026-07-22 @ 21:15:26Points: 134Comments: 38

Fairphone 6 wide camera experimental Linux support

2026-07-22 @ 20:16:01Points: 145Comments: 52

John C. Dvorak has died

2026-07-22 @ 19:22:19Points: 855Comments: 288

Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong

2026-07-22 @ 17:56:29Points: 168Comments: 37

A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on most benchmarks.

- ChartQA: 15-20%

- LibriSpeech: 25-30%

- MMBench, GigaSpeech, MMAU: 30-35%

- MMLU-Pro: 45-55%

We were always frustrated by the routing signals hybrid apps rely on: asking the model to rate itself in text (unreliable, and you're parsing prose), or token entropy heuristics (barely better than a coin flip in our tests). So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.

SO we extended the model with a 68k params probe layer (LayerNorm, low-rank projection, attention pooling, small MLP head) reads one intermediate layer during decoding and predicts p(wrong); confidence = 1 - p(wrong), returned as structured data, never parsed out of the answer text.

Across 12 hold-out benchmarks spanning text, vision and audio, the probe averages 0.814 AUROC vs 0.549 for token entropy. The result that convinced us this is real: the probe was trained on zero audio data, yet scores 0.79-0.88 AUROC on four audio benchmarks where entropy is near-random or worse (0.32-0.52). It's reading a modality-independent correctness signal from the hidden state, not memorizing patterns from its training data.

We published all weights on HuggingFace and provide copy-pase codes to run it on Transformers, MLX, Llama.cpp or Cactus. With Ollama, vLLM, SGLang etc in the works. For llama.cpp we ship a patch series you compile in once (upstreaming is planned). The code is MIT licensed; Gemma model use remains subject to the Gemma terms.

GitHub: https://github.com/cactus-compute/cactus-hybrid

Weights: https://huggingface.co/collections/Cactus-Compute/cactus-hyb...

Some caveats:

- The probe scores single-sequence decoding only, up to the first 1024 generated tokens.

- Handoff works best when routing per task in a multi-step process, not per step.

- Hierarchical routing is still in the works: try on-device, then DeepSeek v4 Flash, before Fable/GPT5.5/Gemini/Muse/Grok.

- The technique is boutique for each model, we will share each weights as they roll out.

These issues are currently being tackled at Cactus and updated weights will be shipped directly into the HuggingFace collection and GitHub repository straight up. Please let us know your thoughts, it helps us find ways to improve the design progressively.

Thanks a million!

Everyone should know SIMD

2026-07-22 @ 17:48:18Points: 544Comments: 197

Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample

2026-07-22 @ 17:30:40Points: 1016Comments: 583

A digestion of the Jacobian conjecture counterexample - https://news.ycombinator.com/item?id=48998362 - July 2026 (133 comments)

Claude Fable produced a counterexample to the Jacobian Conjecture - https://news.ycombinator.com/item?id=48973869 - July 2026 (508 comments)

GigaToken: ~1000x faster Language model tokenization

2026-07-22 @ 17:20:38Points: 576Comments: 115

Are AI labs pelicanmaxxing?

2026-07-22 @ 17:17:54Points: 622Comments: 236

Making

2026-07-22 @ 15:33:48Points: 413Comments: 168

Show HN: Bento - An entire PowerPoint in one HTML file (edit+view+data+collab)

2026-07-22 @ 15:19:23Points: 936Comments: 214

To avoid this loop, I ended up creating Bento, a single HTML file with everything you need in a slide tool including animations and shared editing. There's no install or cloud login, everything works offline. The default deck is around 560 KB and it doesn't need to fetch anything once you got it.

Open it in a browser and then you can edit, present, print and save. Share it via email or via Airdrop and all they need is a browser to edit, present and also do live collab on the slides. Drop it in to Claude or ChatGPT to transform existing pptx files into Bento slides. There is no cloud involved, only an encrypted blind relay to allow for shared editing. The relay doesn't see any of the data.

Check it out at https://bento.page/slides/ which takes you straight to the editor.

Go to https://bento.page/guestbook/ to try out the live guestbook to experience share editing / collab.

There is also a gallery with some sample decks on the website - https://bento.page/

All the code is MIT licensed and you can find it here - https://github.com/nyblnet/bento . I used reveal.js with several other libraries (including some homegrown ones), and Claude Code.

Quality non-fiction books are the antithesis of AI slop

2026-07-22 @ 14:18:05Points: 449Comments: 220

Businesses with ugly AI menu redesigns

2026-07-22 @ 12:49:45Points: 346Comments: 268

The startup's Postgres survival guide

2026-07-22 @ 12:36:08Points: 471Comments: 206

The Unity CLI: manage Unity from your terminal

2026-07-21 @ 17:51:57Points: 47Comments: 14

git's –end-of-options Flag

2026-07-21 @ 13:13:47Points: 175Comments: 111

Escape IntelliJ: Scala and Kotlin LSPs on Emacs Eglot

2026-07-21 @ 10:04:20Points: 115Comments: 93

ANSI escape injection in MCP servers: Hidden from humans, visible to AI

2026-07-21 @ 07:01:24Points: 41Comments: 22

Making ASCII Art in Vim

2026-07-20 @ 18:42:07Points: 102Comments: 12

Frequently Asked Questions on Expertise

2026-07-20 @ 14:00:29Points: 27Comments: 2

Scanning for Pangram Errors

2026-07-17 @ 09:29:52Points: 39Comments: 21

Test-time training 3D reconstruction

2026-07-16 @ 00:11:55Points: 21Comments: 1

Nobody knows what a used GPU cluster is worth

2026-07-15 @ 06:50:49Points: 265Comments: 239

Archives

2026

2025

2024

2023

2022