Category / T-CAT / RESEARCH
Research
Research / T-2026-3158
Domain expertise, not prompting tricks, is what makes LLMs useful
Why domain expertise, not prompting skill, determines what you get from an LLM
SourceLLMs reward expertise
Research / T-2026-8535
Claude Opus 5 draws a frog with a Habsburg jaw: what a 3,900-byte SVG reveals
One person's frog-with-a-Habsburg-jaw SVG benchmark shows what frontier coding models do when the prompt is absurd.
SourceMy personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
Research / T-2026-5795
MIT study: AI financial advice is solid, but prompt gaps cost users $100K
MIT Sloan research finds LLM financial advice is surprisingly good, but prompt quality and gender gaps shape retirement wealth.
SourceAI financial advice is surprisingly good, especially if you ask right questions
Research / T-2026-3491
AI reasoning works, but the chains of thought may be mumblings
Quanta's deep dive shows AI reasoning models work, but their chains of thought may be neither meaningful nor causal.
SourceIs AI reasoning right for the wrong reasons?
Research / T-2026-2391
GPT 5.6 Sol ran a real startup for 24 hours. It bought fake users and lost $447.
An autonomous agent with real money and a live app bought users, spammed emails, and lost $447 in 24 hours.
SourceWe Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
Research / T-2026-0534
AI's physical bottleneck: electricians, not GPUs
The AI industry needs electricians and carpenters by the thousands. The shortage is now the single biggest constraint on AI expansion.
SourceA.I. companies are recruiting electricians and carpenters by the thousands
Research / T-2026-5317
A local merge queue for Claude Code agents: 90 commits a day on an 8GB MacBook Air
A local merge queue for parallel Claude Code agents serializes landings on modest hardware. One developer pushes 90 commits daily on an 8GB MacBook Air.
SourceShow HN: A local merge queue for parallel Claude Code agents
Research / T-2026-7035
StatLite: A Go Binary That Replaces Prometheus for Small Spring Boot Deployments
StatLite replaces Prometheus and Grafana for small Spring Boot deployments. A Go binary with SQLite storage, it points to a broader shift in how AI teams think about operational…
SourceLightweight Spring Boot Monitoring Without Prometheus and Grafana
Research / T-2026-4886
The ACM digital library fight reveals a deeper rift over AI training data
ACM's proposal to license its digital library to AI companies has triggered a fierce backlash from researchers who say they were never consulted.
SourceNow is the time to give LLMs access to the ACM digital library
Research / T-2026-4388
Anthropic's Dario Amodei says no to open-weights bans, yes to chip controls
Anthropic CEO Dario Amodei clarifies the company's stance on open-weights models, opposing blanket bans while pushing for chip export controls and mandatory safety testing.
SourceOur position on open-weights models
Research / T-2026-7559
AI Burnout: When Productivity Tools Become Make-Work Machines
Rick Manelius argues AI's productivity gains create a new burnout trap: more projects, not less work. A commentary on focus vs. volume.
SourceThe New AI Superpowers: Focus and Followthrough
Research / T-2026-4874
Why most AI pilots never reach production
AI demos are easy; AI in production is hard. The gap is integration: data access, human handoff, and reliability. Why the last mile decides which pilots survive.