Hacker News

Latest

The AI Aesthetic

2026-07-30 @ 23:22:16Points: 125Comments: 65

Investigating three real-world incidents in our cybersecurity evaluations

2026-07-30 @ 23:00:51Points: 98Comments: 87

I flagged two research papers for fake authors and both were accepted as orals

2026-07-30 @ 22:33:11Points: 96Comments: 36

Rune 1.1: adds Python, an Emacs editor, a symbol index and is now free

2026-07-30 @ 21:47:31Points: 51Comments: 15

Saber-toothed cats became inbred–and struggled to move–before they went extinct

2026-07-30 @ 21:31:03Points: 27Comments: 7

Agent Skill to Force Docs in ASD-STE100 Simplified Technical English

2026-07-30 @ 19:34:27Points: 212Comments: 78

UEFA and its national associations will not participate in FIFA competitions

2026-07-30 @ 18:40:52Points: 804Comments: 437

Making Postgres queues scale

2026-07-30 @ 18:39:32Points: 104Comments: 26

Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

2026-07-30 @ 18:13:06Points: 85Comments: 61

http://playground.ctgt.ai/

I will now dive in to the motivation, methodology and detailed results for those interested. The hard part of measuring this phenomena is isolating whether a model is reluctant to talk about sensitive things generally vs. a particular country's sensitive things. So we made 152 matched pairs where one prompt asked about a Chinese concept, and the other asked about a non-Chinese version of that concept. For example, the Great Leap Forward vs. the Holodomor. These were scored 0-100 by four LLM judges (Grok 4.20, Gemini 3.5 Flash, GPT-5 mini, Claude Sonnet 4.6), validated against 96 human scores at r=0.948. OpenRouter blocked some of these so we hosted the weights ourselves.

The teacher's gap on the core political set of pairs was +45.45 points, ~7 standard deviations from chance, and every distilled student was within 1 point of its base. Subliminal learning literature says this is expected when the initializations are not shared between teacher and student, which is true here. The distillation data also did not contain any China-sensitive content. The contribution here was to release the evaluation framework (LineageEval: https://github.com/CTGT-Inc/lineage-eval/) to elevate the discussion around this topic in DC and beyond. We are an interpretability lab working on high risk and regulated applications of AI, so we hear a lot of vagaries aimed at the supposed dangers of distilling Chinese models on American bases. We believe these conversations should be based on open, auditable frameworks and not feelings. We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next.

The distillation method was an evolution of HINT-SD where we inject a hint at the specific point the model makes a mistake in its reasoning. Then we train on the corrected continuation with reverse KL over the next 100 toks of the rollout. As mentioned above 120B itself was efficacious as a teacher, and we ended up shipping this version. The self-distilled 120B scores 83.61% on FinanceReasoning, above Kimi K3 (81.93%) and Inkling (65.13%). Ours finishes 98.7% of problems in budget; the larger models truncate (90.76% and 71.01%) which score as incorrect. At 100k tokens big models gain (Kimi 89.92%). So for a finance task at a constrained (perhaps more realistic) budget a 120B on one H100 at ~$0.00026/query outpaced models running 62-160x more per query.

We put out the 20B finance model as open weights (64.71% to 74.79% at 8k on FinanceReasoning, 23% lower cost/query, runs on one 80GB GPU), the 120B in a playground with teacher and students side by side (a few queries, no auth), and LineageEval with all prompts, controls, rubric, and code.

We are curious to hear experiences from those working with distilled Chinese models in prod, or if you have thoughts on improvements to LineageEval.

https://huggingface.co/ctgt-inc/gpt-oss-20b-finance

https://playground.ctgt.ai/

https://github.com/CTGT-Inc/lineage-eval/

https://www.ctgt.ai/research/distillation-censorship-transfe...

CodePen 2.0

2026-07-30 @ 17:52:51Points: 133Comments: 41

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

2026-07-30 @ 17:31:07Points: 300Comments: 186

Advancing the price-performance frontier with GPT‑5.6

2026-07-30 @ 17:15:51Points: 502Comments: 334

Read this before you buy that TV streaming stick

2026-07-30 @ 17:04:53Points: 579Comments: 342

Rise Reforming (YC S26) Is Hiring

2026-07-30 @ 17:00:21Points: 1

Stacked PRs are now live on GitHub

2026-07-30 @ 16:26:16Points: 480Comments: 166

Physicists Solve a Muon Mystery. Now, Old Results Don't Add Up

2026-07-30 @ 15:22:46Points: 182Comments: 94

Gemini Robotics 2 brings whole body intelligence to robots

2026-07-30 @ 15:15:48Points: 481Comments: 396

The Economic Benefit of Refactoring

2026-07-30 @ 15:10:27Points: 197Comments: 81

Hacker Public Radio

2026-07-30 @ 14:29:27Points: 138Comments: 25

The lost civic life of movie rental stores

2026-07-30 @ 14:11:42Points: 126Comments: 186

Upper stage impacting the moon on 2026 August 5

2026-07-30 @ 13:21:34Points: 182Comments: 47

Why is everyone trying to build a solid-state battery?

2026-07-30 @ 12:38:51Points: 175Comments: 208

GCC steering committee announces AI policy

2026-07-30 @ 11:45:44Points: 243Comments: 280

Google will expand age checks on Android worldwide till the end of the year

2026-07-30 @ 10:13:46Points: 355Comments: 437

Show HN: Kedge – Full-stack cloud with forkable VM snapshots and global SQLite

2026-07-29 @ 16:15:57Points: 93Comments: 20

I helped build Fly.io for 4 years and shared enthusiasm for the founders' vision of a 'global Heroku'. While there, I wrote "The Serverless Server" (https://fly.io/blog/the-serverless-server/) as a study of Lambda and a sketch of a modern serverless product built around lightweight VMs. That essay was the initial inspiration for Kedge.

Kedge has a fast VM orchestrator that can create code sandboxes or scale service instances in 3ms, using a combination of forkable VM snapshots and a tree of warm pools (Linux kernel -> base runtime -> app). VMs are memory-dense thanks to shared copy-on-write memory pages. You can run lightweight CGI-style functions, public OCI images, or source code for BuildKit to compile and deploy.

Kedge's global control plane sits on an eventually-consistent SQLite database. Taking inspiration from Corrosion and Litestream, I built a local-first, multi-writer CRDT-based replication system backed by object storage, and just recently made it open source (https://github.com/wjordan/syzy).

You can also use a SQLite client to query `/shared.db` from any instance for a build-in replicated database in your app. This lets Kedge autoscale services close to your users while each instance queries its local replica for eventually-consistent data, with no need to micro-manage instance or volume placement. (There's also a /shared/ filesystem adapter for convenience.)

Kedge can even use this same database for stateful, server-rendered HTML apps. Data attributes bind forms, buttons, and values to records in the app database, Kedge compiles the schema and operations at deploy-time, and then queries the local data to serve requests. As a demo, I made a Hacker News clone with story submission, votes, comments and auth in about 60 lines of Markdown, plus CSS (https://kedge.dev/docs/html-apps#kedger-news).

I've just started collecting public feedback, so please let me know what you think! I'm particularly interested in feedback on the stateful HTML app model, which is the newest (and most ambitious) piece. The preview is currently running in 11 regions for you to kick the tires. There's no billing yet, so the pricing page is an estimate. Thanks for taking a look!

Memo-1: A 6502 computer built from scratch, using a Minitel as its terminal

2026-07-28 @ 13:56:27Points: 56Comments: 6

The American Grilled Cheese Sandwich Essay (2024)

2026-07-27 @ 17:37:32Points: 18Comments: 4

Bad Apple but It's Traceroute

2026-07-27 @ 15:48:09Points: 64Comments: 12

Destroying a Community with a Gigantic "Clogged Vacuum Cleaner"

2026-07-27 @ 12:33:50Points: 62Comments: 38

2x, not 10x: coding with LLMs in 2026

2026-07-25 @ 14:27:05Points: 214Comments: 170

Archives

2026

2025

2024

2023

2022