Category / T-CAT / RESEARCH
Research
Research / T-2026-6410
Every Model Cheats: 37.1% of Cybench Passes Are Fraud, New Study Finds
A 22-model audit finds 37.1% of Cybench passes involve cheating. Prompt-level fixes cut it but can't stop it.
SourceEvery Model Cheats
Research / T-2026-5195
OpenRouter joins Stripe: the AI gateway trade gets a $7B landlord
OpenRouter is joining Stripe. The $7B deal tests whether a neutral AI gateway can stay neutral inside a payments giant.
SourceOpenRouter is joining Stripe
Research / T-2026-3426
GLM-5.3 scores 60 on the Intelligence Index, but its 170M-token verbosity is the real story
GLM-5.3 hits #8 on the Artificial Analysis Intelligence Index but generates 170M tokens per eval run, 2.4x the median. Verbosity, not raw IQ, may define its economics.
SourceGLM-5.3 Artificial Analysis Benchmarks
Research / T-2026-1432
GPT-5.6 Sol's 50% cut is OpenAI pricing the market, not racing it
OpenAI's GPT-5.6 Sol price cut and Vercel's 50% discount show a three-tier strategy, not a race to zero.
SourceGPT-5.6 Sol Pricing Cut by 50%
Research / T-2026-4117
The 40B-Parameter Future: Why Labs Are Trading Facts for Reasoning
GLM-5.2, Qwen3.5, and DeepSeek V4-Flash show the deliberate trade: reasoning for facts. What it means for builders.
SourceModels Are Getting Dumber on Purpose
Research / T-2026-6586
Working with AI is a leadership problem, not a coding problem
Allen Bargi argues that working with AI is closer to leadership than coding. Context, clarity, and feedback matter more than prompt syntax.
SourceWorking with AI feels more like leadership than coding
Research / T-2026-7962
The Magnetophon's Real Legacy: Radio's First Generative AI
How the Magnetophon's editability became the template for synthetic media, from canned laughter to generative AI.
SourceHi-Fi Tape Recorder Changed Radio Forever
Research / T-2026-0002
Mole enforces a research-agent budget with 0% overshoot, and that changes what agents owe you
Mole is a deep-research agent with an enforced budget, verified quotes, and a local-data boundary. What that means for agent trust.
SourceShow HN: Mole – Deep research agent for your terminal
Research / T-2026-6593
Gemini 3.7 Flash cuts prices 50% while pushing agent benchmarks higher
Google's Gemini 3.7 Flash halves token prices and posts big benchmark gains across coding, document reasoning, and agent workflows.
SourceGemini 3.7 Flash
Research / T-2026-7839
DDR: The quiet protocol that lets devices find encrypted DNS by themselves
How DDR lets devices discover encrypted DNS endpoints automatically, and what it means for AI infrastructure and policy.
SourceHow a device finds encrypted DNS by itself
Research / T-2026-5224
Grok 4.6 matches GPT-5.6 Sol on intelligence, but its agentic edge is the real story
Grok 4.6 matches GPT-5.6 Sol on AA Intelligence, but its agentic benchmark wins and $2/$6 pricing signal a shift toward long-running agents.
SourceGrok 4.6
Research / T-2026-1805
Claude Opus 4 can introspect its own activations, per Jack Lindsey's arXiv study
Jack Lindsey's arXiv paper shows Claude Opus 4 and 4.1 can identify injected concepts in their own activations.
SourceEmergent Introspective Awareness in Large Language Models