from/prod
← All companies

THE COMPANY INDEX TRACKED BLOG

Stack Overflow

Ideas, decisions, and lessons from the team.

stackoverflow.blog (opens on the source site)LinkedIn X
57Posts tracked
yesterdayLatest publication
4.8Posts / month over the last 12 months

Latest writing

20 of 57 posts

A Treatise on Model Oriented Programming Languages (opens on the source site)

Thirteen years ago I developed to my knowledge the first and currently the only fully featured model oriented programming language and IDE. I utilized this technology (Mo+) to great effect on my own and workplace enterprise projects, but failed to sell the ideas to a wider audience. Six years ago I left my software engineering career in favor of wielding an ax and building Viking and Anglo Saxon ships in Norway, UK and other future places in Scandinavia. Even so, I still think about model oriented programming from time to time and its potential. The purpose of this article is not about the…

Read at the source

Taking a look under your agent’s hood (opens on the source site)

Ryan is joined by Yanbing Li, Chief Product Officer at Datadog, to talk about applying observability to non-deterministic AI agents, blurring the boundaries between software development and production workflows, and navigating emerging challenges in AI security and tokenomics.

Read at the source

Part 4: Safety and governance for LLM systems: guardrails, PII, audit, and memory (opens on the source site)

The level where an LLM system stops being a demo and earns the right to touch real data and real decisions: layered guardrails that fail closed, PII handled at the boundary, an immutable audit trail, and scoped memory. By the time an LLM system is making decisions that matter, “it usually works” is no longer the bar. This is Level 4 of the maturity model — safety and governance — and it’s where four disciplines that teams tend to bolt on late have to be designed in instead. They share one idea: don’t trust a single point to do the right thing. Layer independent guardrails so a miss at one is…

Read at the source

Part 3: Knowing when your agent doesn’t know: the confidence layer (opens on the source site)

The most important number an agent produces isn’t its answer — it’s how sure it is. Compose that number from independent signals, check it’s calibrated, grade the high-stakes calls with a second model, and route the rest to humans well. This is Level 3 of a six-level maturity model for running LLM systems in production. Levels 1 and 2 got you to where the system works and you can see it working. Level 3 is confidence: the system acts on its own only when its calibrated confidence is high, grades the decisions that matter with an independent judge, and routes everything it’s unsure about to a…

Read at the source

Part 2: Evals as a deployment gate — and how to know when they drift (opens on the source site)

If you can deploy a prompt change without an eval failing the build, you don’t have evals — you have a notebook. And once the gate is green, the slow leaks are still coming for you. Here’s the gate, the baseline, and the shadow-eval loop that catch both. This is Level 2 of the maturity model: evaluation. The principle is short — you don’t ship on hope, you ship on a gate, and then you watch for drift afterward. A gate protects the moment of deploy. Drift detection protects the weeks in between. You need both, and they’re built from different machinery. Most teams “do evals” the way they once…

Read at the source

Part 1: Make your AI agents boring: the determinism layer (opens on the source site)

The trick to putting LLM agents in high-stakes systems isn’t a smarter model — it’s containing the model to one node so the rest of the system is ordinary, testable code. Here are the structural moves, with the contracts and types to implement them. Demos love autonomous agents that loop, call tools, and “figure it out.” Production hates them. The moment an agent’s behavior depends on which path the model wandered down today, you can’t test it, can’t audit it, and can’t let it touch anything that matters. In a regulated or high-consequence system — money movement, healthcare, infrastructure —…

Read at the source

Implementing a Modular Master-Agent Telemetry & Diagnostic Framework in Python: Prime-Sentinel Command (PSC) (opens on the source site)

Tags: python architecture oop design-patterns distributed-systems When engineering distributed monitoring agents or designing low-latency health-checking pipelines, separating centralized governance from autonomous edge execution is essential. I designed the Prime-Sentinel Command (PSC) architecture as an object-oriented master-agent pattern to coordinate edge diagnostic nodes (Sentinels) via a centralized orchestrator (Prime). Below is an architectural walkthrough and minimal reference implementation for engineers looking to build similar decoupled telemetry collectors. Many diagnostic…

Read at the source

Tales from the 2026 Developer Survey results (opens on the source site)

Ryan chats with Erin Yepis, Senior Analyst at Stack Overflow, about the results from this year’s Annual Developer Survey, including the overwhelming daily usage of AI coding assistants despite lingering developer trust issues, the critical role of well-organized documentation in providing verifiable context to mitigate AI hallucinations, and the evolving ways developers are shifting away from active community posting in favor of passive knowledge consumption.

Read at the source

Getting ready for 2026 results: A look back on Developer Survey findings (opens on the source site)

We are on the precipice of brand new results from the 2026 Developer Survey. We have 15 years of data and insights from this survey at our fingertips that shows more than just what one year’s collection of insights will explore. A year-over-year look at the Developer Survey offers the long view of what is changing in the broader technology ecosystem. This analysis compares the 2024 and 2025 Developer Survey data across three connected stories: the evolution of AI, humans at work, and demographics and community. The survey results reveal what the headline news cannot: how people actually fold…

Read at the source

Is Your “Human-in-the-Loop” Actually Slowing You Down? Here’s What We Learned (opens on the source site)

In the rush to adopt AI and automation, many teams implement human-in-the-loop (HITL) frameworks. They believe that involving a person in the process solves the problems with reliability, quality, and trust. But as we’ve learned from real engineering workflows and integrations, the story isn’t that easy. In some contexts, humans-in-the-loop do improve outcomes, but in others, they can unintentionally become bottlenecks that limit speed, scalability, and innovation. In this post, we’ll analyze when human-in-the-loop is truly valuable, when it slows systems down, and how to strike the right…

Read at the source

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.