Hacker News

Latest

Half of Europe's towns and villages have fewer residents than 60 years ago

2026-08-11 @ 05:46:16Points: 43Comments: 65

DeepSeek: Reverse Engineering an AI Assistant by Interviewing Itself

2026-08-11 @ 05:30:24Points: 20Comments: 5

Show HN: Mcptoon – MCP CLI client that cuts tool discovery tokens by 97%

2026-08-11 @ 05:26:58Points: 27Comments: 15

Microsoft Responds to Outcry After Quiet Enterprise Install of Beta 'Photos' App

2026-08-11 @ 04:18:37Points: 41Comments: 27

Updated GPG Key for Signing Firefox and Thunderbird Releases

2026-08-11 @ 04:04:36Points: 32Comments: 11

Why My Father Is Wrong: A Defense of Guitar Hero

2026-08-11 @ 03:52:39Points: 41Comments: 19

Hyperspace

2026-08-11 @ 02:04:35Points: 60Comments: 42

Recycle – Floppydisks

2026-08-11 @ 02:01:01Points: 51Comments: 19

H3-metal – Native MiniMax-H3 inference for Apple Silicon

2026-08-11 @ 01:22:09Points: 205Comments: 26

Chicken Scheme 6.0

2026-08-11 @ 00:24:15Points: 155Comments: 17

Show HN: Scroll through all 43252003274489856000 Rubik's Cube states

2026-08-10 @ 23:16:25Points: 166Comments: 52

World Train Map – 1247 train routes around the world

2026-08-10 @ 22:42:14Points: 115Comments: 41

Confessions of a Long-Distance Sailor

2026-08-10 @ 20:52:55Points: 98Comments: 28

Rust SIMD on the GPU

2026-08-10 @ 18:12:49Points: 170Comments: 83

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

2026-08-10 @ 17:22:07Points: 294Comments: 108

Henry from Cactus here!

We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2.

The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges 300-700 on sub-$200 phones such as the Samsung A-Series.

On the tool call and mobile device use benchmarks, Needle 2 trades wins with closest small models like LFM2.5 230M and Apple Foundation Model, at 5x to 70x smaller, both at f16 vs Needle 2 at 2bit. Needle is based on Simple Attention Networks from our paper (https://arxiv.org/abs/2607.18363).

Edge AI has lately meant Macs and PCs, but that is just 1.5 billion of over 21 billion connected IoT devices in the world today, and in emerging markets most phones ship under $200, no NPU, cheap GPUs. These include budget phones, Raspberry Pis, microcontrollers, wearables, small robots like Reachy Mini, and connected home devices.

A conventional transformer of Needle's width and depth spends 164 MFLOPs per token, and even one squeezed down to Needle's parameter count spends 87, Needle spends 70. Even on a high-end phone, an always-on assistant lives inside a power budget; every MFLOP is milliwatt-hours, and Needle spends 7x to 85x fewer of them per token than the smallest performant LLMs. More about the architecture in the link.

When we structure intelligence for consumer devices as functions with typed parameters, the only hard part is mapping a messy sentence onto them; which function, with which values. Our research found that when framed that way, the problem needs no world knowledge and no open-ended prose, which is why 45M parameters suffice.

Needle 2 expands to structured extraction where the schema can be passed in-place of tools and the model returns structured output. You can use Needle as a text-classification model with an enum field, as a summarization model by providing a schema that extracts key fields, everything but free-range decode.

Every product has its own tool vocabulary and fine-tuning needle helps it achieve frontier-level performance on custom tasks, so using the python package (https://github.com/cactus-compute/needle), Needle can be fine-tuned Needle on a Mac/PC in minutes to a few hours, with automated data-generation pipeline, just pass a couple samples.

Nonetheless, every response carries a learned confidence score based our Cactus Hybrid technique. If above your threshold, act, below it, escalate to the cloud or bigger model. Combining Needle 2 with a private DeepSeek-v4-Flash deployment works particularly well for enterprise-level tasks at barely any cost, we can help with this setup.

We have put a lot of thoughts into Needle 2 but might still be missing quite a lot, please use the playground in the provided link to test Needle and share your thoughts, always appreciated!

Launch HN: Stoa Markets (YC S26) – A Marketplace for GPUs and AI Servers

2026-08-10 @ 16:35:27Points: 75Comments: 50

https://www.stoaexchange.com), a marketplace for new and used GPUs and AI servers.

GPUs are the collateral in the data center buildout. Today, financing terms mostly depend on the offtaker, meaning the company that has committed to use the compute.

If that company is a hyperscaler, the financing can look investment grade. If it’s a smaller cloud or startup, terms get expensive fast, even with the same hardware as collateral.

The lender’s problem is pretty reasonable. If the borrower defaults and we need to sell these servers, what can we actually get for them? There isn’t a good answer today.

We started brokering GPU deals to understand why. It was much more manual than we expected. The hardware is still traded through phone calls, forwarded spreadsheets and long email threads.

One week, a seller quoted us $200k for a server node and another quoted $240k for what looked like the same thing. Neither was necessarily wrong. They had different information and could only see their own corner of the market.

Before we could compare the quotes, we had to sort out the configuration, condition, warranty, location and delivery terms. It’s the same information Kelley Blue Book attaches to a used-car price through the year, trim, mileage and condition. “An H100 server” isn’t enough information to know what something is worth, just as “a used BMW” isn’t.

This is also just a bad way to buy or sell hardware. A buyer looking for the best price shouldn’t have to contact several brokers and dealers separately, repeat the same request and then untangle a pile of different quotes. Sellers shouldn’t have to search for demand one buyer at a time. A market of this size deserves better liquidity.

Stoa puts the request into one format and sends it to dealers that have gone through know-your-business (KYB) checks. We verify the company, who owns it and who is allowed to trade for it. Before the request goes out, the buyer confirms the exact configuration, quantity, condition, warranty, location, delivery terms and what will be checked during inspection. Dealers return firm quotes against that same request without seeing each other’s bids. Once a quote is accepted, payment, shipping, delivery and inspection are then tracked through settlement. We don’t take possession of the hardware.

We got more than $300M in requests for quotes (RFQs) during our first month.

The immediate goal is to make buying and selling this hardware less painful. As trades build up, they also leave lenders with actual resale evidence instead of list prices and one off appraisals.

We knew from the beginning that this couldn’t be a software only marketplace. GPU trading runs on relationships, and inventory isn’t shown to just anyone. Dealers need to trust the people bringing them clients, and clients need to trust that quotes will actually turn into trades. We built those relationships over time by brokering deals ourselves. Stoa gives people a cleaner way to trade, from the first RFQ through settlement, but it doesn’t replace the trust underneath. Those relationships, and the history of who actually follows through, are a big part of our process.

We’ve known each other for more than ten years. We have founded companies, traded interest rate derivatives, built trading and pricing systems for oil and gas. We learned GPU trading by doing the deals ourselves, and Stoa grew out of the problems we kept running into.

We charge a tiered fee on completed trades, with lower fees at higher volumes. It’s free to sign up at https://www.stoaexchange.com/signup. If you’ve bought, sold, financed or had to liquidate GPUs, would be great to hear your take!

What's the best programming language for coding agents?

2026-08-10 @ 16:28:58Points: 150Comments: 100

Which programming languages are most token-efficient? - https://news.ycombinator.com/item?id=46582728 - Jan 2026 (91 comments)

Exploiting System Management Mode with a very long interrupt

2026-08-10 @ 16:03:14Points: 152Comments: 55

Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models

2026-08-10 @ 14:06:22Points: 477Comments: 436

Humanising LLM Outputs Is Dumb

2026-08-10 @ 13:35:40Points: 196Comments: 128

Squeak 6.1

2026-08-10 @ 12:15:53Points: 256Comments: 125

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

2026-08-10 @ 10:10:02Points: 1098Comments: 600

Choral: Choreographic Programming for Java

2026-08-07 @ 19:21:42Points: 21Comments: 3

Publishing Schematics Before “Open Source” Was a Word

2026-08-07 @ 15:59:08Points: 82Comments: 18

Stowaway – Take the window seat on any plane or satellite overhead

2026-08-07 @ 13:15:10Points: 228Comments: 28

Sonic Pi v5

2026-08-07 @ 10:26:34Points: 359Comments: 87

Once More in Triple Time

2026-08-06 @ 19:08:36Points: 3Comments: 0

The “mechanical miracle” that ruined Mark Twain’s life

2026-08-05 @ 15:22:23Points: 118Comments: 55

The Story of Mac: A Just-So Story

2026-08-05 @ 05:24:14Points: 24Comments: 4

LFM2.5 2.6B model competitive with 4x larger models

2026-08-04 @ 18:47:25Points: 45Comments: 12

Archives

2026

2025

2024

2023

2022