Hacker News

Latest

Google Search Is Dying. What Comes Next Is Worse

2026-08-10 @ 22:36:30Points: 50Comments: 35

Amazon backs power plant that may become top source of US climate pollution

2026-08-10 @ 21:26:16Points: 138Comments: 86

Confessions of a Long-Distance Sailor

2026-08-10 @ 20:52:55Points: 45Comments: 7

Stop Killing Games: It's time to sue Sony, join us

2026-08-10 @ 20:47:13Points: 119Comments: 48

Illinois Just Passed a Law That Puts Linux on the Hook for Age Verification

2026-08-10 @ 20:20:06Points: 270Comments: 339

Rust SIMD on the GPU

2026-08-10 @ 18:12:49Points: 110Comments: 54

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

2026-08-10 @ 17:22:07Points: 114Comments: 61

Henry from Cactus here!

We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2.

The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges 300-700 on sub-$200 phones such as the Samsung A-Series.

On the tool call and mobile device use benchmarks, Needle 2 trades wins with closest small models like LFM2.5 230M and Apple Foundation Model, at 5x to 70x smaller, both at f16 vs Needle 2 at 2bit. Needle is based on Simple Attention Networks from our paper (https://arxiv.org/abs/2607.18363).

Edge AI has lately meant Macs and PCs, but that is just 1.5 billion of over 21 billion connected IoT devices in the world today, and in emerging markets most phones ship under $200, no NPU, cheap GPUs. These include budget phones, Raspberry Pis, microcontrollers, wearables, small robots like Reachy Mini, and connected home devices.

A conventional transformer of Needle's width and depth spends 164 MFLOPs per token, and even one squeezed down to Needle's parameter count spends 87, Needle spends 70. Even on a high-end phone, an always-on assistant lives inside a power budget; every MFLOP is milliwatt-hours, and Needle spends 7x to 85x fewer of them per token than the smallest performant LLMs. More about the architecture in the link.

When we structure intelligence for consumer devices as functions with typed parameters, the only hard part is mapping a messy sentence onto them; which function, with which values. Our research found that when framed that way, the problem needs no world knowledge and no open-ended prose, which is why 45M parameters suffice.

Needle 2 expands to structured extraction where the schema can be passed in-place of tools and the model returns structured output. You can use Needle as a text-classification model with an enum field, as a summarization model by providing a schema that extracts key fields, everything but free-range decode.

Every product has its own tool vocabulary and fine-tuning needle helps it achieve frontier-level performance on custom tasks, so using the python package (https://github.com/cactus-compute/needle), Needle can be fine-tuned Needle on a Mac/PC in minutes to a few hours, with automated data-generation pipeline, just pass a couple samples.

Nonetheless, every response carries a learned confidence score based our Cactus Hybrid technique. If above your threshold, act, below it, escalate to the cloud or bigger model. Combining Needle 2 with a private DeepSeek-v4-Flash deployment works particularly well for enterprise-level tasks at barely any cost, we can help with this setup.

We have put a lot of thoughts into Needle 2 but might still be missing quite a lot, please use the playground in the provided link to test Needle and share your thoughts, always appreciated!

Launch HN: Stoa Markets (YC S26) – A Marketplace for GPUs and AI Servers

2026-08-10 @ 16:35:27Points: 62Comments: 39

https://www.stoaexchange.com), a marketplace for new and used GPUs and AI servers.

GPUs are the collateral in the data center buildout. Today, financing terms mostly depend on the offtaker, meaning the company that has committed to use the compute.

If that company is a hyperscaler, the financing can look investment grade. If it’s a smaller cloud or startup, terms get expensive fast, even with the same hardware as collateral.

The lender’s problem is pretty reasonable. If the borrower defaults and we need to sell these servers, what can we actually get for them? There isn’t a good answer today.

We started brokering GPU deals to understand why. It was much more manual than we expected. The hardware is still traded through phone calls, forwarded spreadsheets and long email threads.

One week, a seller quoted us $200k for a server node and another quoted $240k for what looked like the same thing. Neither was necessarily wrong. They had different information and could only see their own corner of the market.

Before we could compare the quotes, we had to sort out the configuration, condition, warranty, location and delivery terms. It’s the same information Kelley Blue Book attaches to a used-car price through the year, trim, mileage and condition. “An H100 server” isn’t enough information to know what something is worth, just as “a used BMW” isn’t.

This is also just a bad way to buy or sell hardware. A buyer looking for the best price shouldn’t have to contact several brokers and dealers separately, repeat the same request and then untangle a pile of different quotes. Sellers shouldn’t have to search for demand one buyer at a time. A market of this size deserves better liquidity.

Stoa puts the request into one format and sends it to dealers that have gone through know-your-business (KYB) checks. We verify the company, who owns it and who is allowed to trade for it. Before the request goes out, the buyer confirms the exact configuration, quantity, condition, warranty, location, delivery terms and what will be checked during inspection. Dealers return firm quotes against that same request without seeing each other’s bids. Once a quote is accepted, payment, shipping, delivery and inspection are then tracked through settlement. We don’t take possession of the hardware.

We got more than $300M in requests for quotes (RFQs) during our first month.

The immediate goal is to make buying and selling this hardware less painful. As trades build up, they also leave lenders with actual resale evidence instead of list prices and one off appraisals.

We knew from the beginning that this couldn’t be a software only marketplace. GPU trading runs on relationships, and inventory isn’t shown to just anyone. Dealers need to trust the people bringing them clients, and clients need to trust that quotes will actually turn into trades. We built those relationships over time by brokering deals ourselves. Stoa gives people a cleaner way to trade, from the first RFQ through settlement, but it doesn’t replace the trust underneath. Those relationships, and the history of who actually follows through, are a big part of our process.

We’ve known each other for more than ten years. We have founded companies, traded interest rate derivatives, built trading and pricing systems for oil and gas. We learned GPU trading by doing the deals ourselves, and Stoa grew out of the problems we kept running into.

We charge a tiered fee on completed trades, with lower fees at higher volumes. It’s free to sign up at https://www.stoaexchange.com/signup. If you’ve bought, sold, financed or had to liquidate GPUs, would be great to hear your take!

Exploiting System Management Mode with a very long interrupt

2026-08-10 @ 16:03:14Points: 116Comments: 38

Extreme 220GHz+Broadband Silicon Capacitor X2SC 0201M 22nF BV11

2026-08-10 @ 15:53:01Points: 58Comments: 23

Magnitude 7.4 Earthquake – 5 km S of San José del Palmar, Colombia

2026-08-10 @ 15:49:48Points: 155Comments: 61

Mars Bar from 1991 found – and it's 20g bigger than today's

2026-08-10 @ 15:32:04Points: 289Comments: 446

Letter to Governor Abbott on responsible AI infrastructure in Texas

2026-08-10 @ 14:38:20Points: 83Comments: 156

Exploring Claude/GPT Knowledge Cutoffs and Pre-Training Timelines

2026-08-10 @ 14:20:41Points: 92Comments: 13

Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models

2026-08-10 @ 14:06:22Points: 337Comments: 357

Humanising LLM Outputs Is Dumb

2026-08-10 @ 13:35:40Points: 140Comments: 81

Mistral Patent for “Code implemented tool calls”

2026-08-10 @ 13:29:12Points: 205Comments: 170

50k Boat Names

2026-08-10 @ 12:58:16Points: 149Comments: 103

Tl;dv: Over 180k meetings left wide open

2026-08-10 @ 12:26:05Points: 522Comments: 173

Squeak 6.1

2026-08-10 @ 12:15:53Points: 211Comments: 109

Tail-call optimization in C is relatively recent (2025)

2026-08-10 @ 11:34:40Points: 114Comments: 110

Parametron: 50s Japanese computer that uses neither transistors nor vacuum tubes

2026-08-10 @ 10:29:14Points: 176Comments: 45

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

2026-08-10 @ 10:10:02Points: 994Comments: 561

Docker Sandboxes – Disposable, isolated sandboxes for AI agents

2026-08-10 @ 06:02:38Points: 623Comments: 348

Ask HN: In your experience, what are sound conventions for e-ink UI development?

2026-08-07 @ 17:29:16Points: 135Comments: 47

I've recently switched to a black-and-white e-ink smartphone with the motivation of withdrawing from the attention economy somewhat and it's a genuinely cool piece of hardware. While the majority of my needs are met by this device there are a few things I'd like to have which don't work terribly well with the e-ink screen. I'm planning to implement a couple of projects to fill these gaps, at the moment I'm planning a Lemmy frontend and an OpenRouter frontend specifically for e-ink. Both are to be browser-based rather than native, to maximise compatibility and because I'm much more familiar with the web than Android development.

I would like to study the principles of sound e-ink UI design before approaching these projects to avoid creating unusable slop, in particular I am not entirely sure how to approach treating the refreshes as a first-class aspect of the design when I can't control them from the browser, and how to apply comprehensible UI conventions when a greyscale, high-contrast display is the target.

Some specific problems I have out of the gate are:

* Streaming LLM output to the screen is basically the worst-case scenario for e-ink, I need to buffer it and paint it in chunks without this becoming horrible to use.

* Ghosting is a serious problem, browsing HN on the device is a particularly obvious example. Ideally I want to avoid scrolling as far as possible and rely on pagination instead, which I feel has the potential to become annoying if not done well.

* Given I must rely exclusively on layout and type to carry the UI, what design languages emphasise these qualities best? My gut says the early Mac OS versions wouldn't be a bad place to start, this seems relevant given the display constraints of the early macs.

I would greatly appreciate any advice on the design and implementation of e-ink UIs from people with practical experience in this area. This is purely to scratch a personal itch, once they're nailed down I'll put them out in the wild under the GPL.

Publishing Schematics Before “Open Source” Was a Word

2026-08-07 @ 15:59:08Points: 45Comments: 12

Stowaway – take the window seat on any plane or satellite overhead

2026-08-07 @ 13:15:10Points: 77Comments: 7

Sonic Pi v5

2026-08-07 @ 10:26:34Points: 289Comments: 76

Show HN: Higher-dimensional lattices unfolded into the 2D plane

2026-08-06 @ 21:52:31Points: 13Comments: 1

Visit the URL to explore what happens. Double-click or press “N” to generate a new experiment. Press “A” to learn more about how it works.

The Psychedelic Toad of the Sonoran Desert

2026-08-04 @ 17:49:29Points: 69Comments: 48

Archives

2026

2025

2024

2023

2022