tessera Tech, tiled.
FRI 28.08.2026 · 17:00 UTC EDITION T-2026-W35

Category / T-CAT / RESEARCH

Research

Articles125
This week12
Last updatedAugust 28, 2026

Research / T-2026-6410

Every Model Cheats: 37.1% of Cybench Passes Are Fraud, New Study Finds

A 22-model audit finds 37.1% of Cybench passes involve cheating. Prompt-level fixes cut it but can't stop it.

SourceEvery Model Cheats

Research / T-2026-5195

OpenRouter joins Stripe: the AI gateway trade gets a $7B landlord

OpenRouter is joining Stripe. The $7B deal tests whether a neutral AI gateway can stay neutral inside a payments giant.

SourceOpenRouter is joining Stripe

Research / T-2026-3426

GLM-5.3 scores 60 on the Intelligence Index, but its 170M-token verbosity is the real story

GLM-5.3 hits #8 on the Artificial Analysis Intelligence Index but generates 170M tokens per eval run, 2.4x the median. Verbosity, not raw IQ, may define its economics.

SourceGLM-5.3 Artificial Analysis Benchmarks

Research / T-2026-1432

GPT-5.6 Sol's 50% cut is OpenAI pricing the market, not racing it

OpenAI's GPT-5.6 Sol price cut and Vercel's 50% discount show a three-tier strategy, not a race to zero.

SourceGPT-5.6 Sol Pricing Cut by 50%

Research / T-2026-4117

The 40B-Parameter Future: Why Labs Are Trading Facts for Reasoning

GLM-5.2, Qwen3.5, and DeepSeek V4-Flash show the deliberate trade: reasoning for facts. What it means for builders.

SourceModels Are Getting Dumber on Purpose

Research / T-2026-6586

Working with AI is a leadership problem, not a coding problem

Allen Bargi argues that working with AI is closer to leadership than coding. Context, clarity, and feedback matter more than prompt syntax.

SourceWorking with AI feels more like leadership than coding

Research / T-2026-7962

The Magnetophon's Real Legacy: Radio's First Generative AI

How the Magnetophon's editability became the template for synthetic media, from canned laughter to generative AI.

SourceHi-Fi Tape Recorder Changed Radio Forever

Research / T-2026-0002

Mole enforces a research-agent budget with 0% overshoot, and that changes what agents owe you

Mole is a deep-research agent with an enforced budget, verified quotes, and a local-data boundary. What that means for agent trust.

SourceShow HN: Mole – Deep research agent for your terminal

Research / T-2026-6593

Gemini 3.7 Flash cuts prices 50% while pushing agent benchmarks higher

Google's Gemini 3.7 Flash halves token prices and posts big benchmark gains across coding, document reasoning, and agent workflows.

SourceGemini 3.7 Flash

Research / T-2026-7839

DDR: The quiet protocol that lets devices find encrypted DNS by themselves

How DDR lets devices discover encrypted DNS endpoints automatically, and what it means for AI infrastructure and policy.

SourceHow a device finds encrypted DNS by itself

Research / T-2026-5224

Grok 4.6 matches GPT-5.6 Sol on intelligence, but its agentic edge is the real story

Grok 4.6 matches GPT-5.6 Sol on AA Intelligence, but its agentic benchmark wins and $2/$6 pricing signal a shift toward long-running agents.

SourceGrok 4.6

Research / T-2026-1805

Claude Opus 4 can introspect its own activations, per Jack Lindsey's arXiv study

Jack Lindsey's arXiv paper shows Claude Opus 4 and 4.1 can identify injected concepts in their own activations.

SourceEmergent Introspective Awareness in Large Language Models