Engineering, AI & what's being built.
Deep-dives, model benchmarks, and field notes from the people shipping the infrastructure.
SWE-bench: What It Actually Measures, and Where It Misleads
SWE-bench Verified is the benchmark every AI coding tool cites. But the score you see quoted often obscures what the test actually measures — and what it does…
The Small Model Renaissance: When Big AI Doesn't Win
The narrative that big always wins is a partial narrative. We map the class of small, task-focused models that quietly took a growing share of production AI…
Zero-Trust in Production: A 2026 Field Guide
Zero-trust is no longer a buzzword — it is the default security architecture for serious organizations. We map what production deployments actually look like…
The archive
After the AI Bubble: An Audit of 2025's Promises
Eighteen months ago the AI conversation was wall-to-wall predictions. We pulled the most-cited public claims about 2026 from late 2024 — AGI timelines, job…
The Real State of MCP: 18 Months in Production
Eighteen months after Anthropic open-sourced the Model Context Protocol, MCP has quietly become the default plug for AI agents. We surveyed the public registry…

GitHub's 10,000-Repo Trojan: The Supply Chain Attack Reshaping Software Security
The discovery of 10,000 GitHub repositories actively distributing Trojan malware marks a critical inflection point in software supply chain security. This…

Lore: The Next-Gen Version Control Paradigm for Petabyte Monorepos & Global Teams
Lore Version Control: A New Paradigm for Petabyte Monorepos & Global Teams Git's Unbearable Weight: When a Standard Becomes an Impediment The reality of modern…
Qwen3.6-Plus: A Leap Forward in Real-World Agents
Qwen3.6-Plus is a significant improvement over its predecessor, offering better performance and adaptability in real-world scenarios.
Microsoft's GUI Strategy: A Critical Analysis
Microsoft's GUI strategy has been criticized for being inconsistent and confusing. But what's behind this criticism, and what does it mean for users?
Choosing Your First AI Infra Stack: A Founder's Field Guide for 2026
An opinionated, no-nonsense guide to assembling your first production AI stack in 2026 — what to pick, what to skip, and what to defer until Series A.
Wii Runs Mac OS X
Discover how one developer managed to port Mac OS X to the Nintendo Wii, and what this means for the world of console hacking. Learn about the challenges and…
The GPUs That Shaped the Industry
From humble beginnings to cutting-edge technology, we explore the GPUs that revolutionized computing.
How I Cut Our Anthropic Bill by 84%: A Prompt Caching Playbook for 2026
Most teams treat Claude's prompt caching like a checkbox. Here's the production tuning playbook from three companies that dropped their bills 70-85% in a month.
The Rise of the Claude Skills and Agent SDK Ecosystem
The Anthropic Agent SDK and Claude Skills ecosystem went from new toy to default in roughly nine months. Here is what they are, why they won, and what to build…

Unlocking the Power of Local AI: How Laptops Are Revolutionizing Artificial Intelligence
The shift towards local AI is transforming the way we interact with artificial intelligence. With the ability to run sophisticated models directly on laptops…

Revolutionizing AI: How Rio's Modular Approach to LLM Integration Is Redefining Industry Standards
Rio's pioneering work in large language model development is set to disrupt the status quo, offering a more accessible and specialized AI solution. By merging…