Thirteen years ago I developed to my knowledge the first and currently the only fully featured model oriented programming language and IDE. I utilized this technology (Mo+) to great effect on my own and workplace enterprise projects, but failed to sell the ideas to a wider audience. Six years ago I left my software engineering career in favor of wielding an ax and building Viking and Anglo Saxon ships in Norway, UK and other future places in Scandinavia. Even so, I still think about model oriented programming from time to time and its potential. The purpose of this article is not about the…
Ryan is joined by Yanbing Li, Chief Product Officer at Datadog, to talk about applying observability to non-deterministic AI agents, blurring the boundaries between software development and production workflows, and navigating emerging challenges in AI security and tokenomics.
A maturity model for taking agents from an impressive demo to a system people can depend on — with a self-assessment and the map to a deep-dive on each layer.
Your service can be 100% up and still quietly approving the wrong things, burning its budget, or failing over into untested quality. Level 5 is the infrastructure that lets you see your decisions, bound your spend, route and fail over between models, kill bad behavior in seconds — and the platform that makes all of it possible.
Agents don't build trust for another reason, structurally worse than the first. It isn't only that the tool keeps changing shape. It's that the feedback loop you would need in order to learn the tool is broken at the point of measurement.
The level where an LLM system stops being a demo and earns the right to touch real data and real decisions: layered guardrails that fail closed, PII handled at the boundary, an immutable audit trail, and scoped memory. By the time an LLM system is making decisions that matter, “it usually works” is no longer the bar. This is Level 4 of the maturity model — safety and governance — and it’s where four disciplines that teams tend to bolt on late have to be designed in instead. They share one idea: don’t trust a single point to do the right thing. Layer independent guardrails so a miss at one is…
The most important number an agent produces isn’t its answer — it’s how sure it is. Compose that number from independent signals, check it’s calibrated, grade the high-stakes calls with a second model, and route the rest to humans well. This is Level 3 of a six-level maturity model for running LLM systems in production. Levels 1 and 2 got you to where the system works and you can see it working. Level 3 is confidence: the system acts on its own only when its calibrated confidence is high, grades the decisions that matter with an independent judge, and routes everything it’s unsure about to a…
If you can deploy a prompt change without an eval failing the build, you don’t have evals — you have a notebook. And once the gate is green, the slow leaks are still coming for you. Here’s the gate, the baseline, and the shadow-eval loop that catch both. This is Level 2 of the maturity model: evaluation. The principle is short — you don’t ship on hope, you ship on a gate, and then you watch for drift afterward. A gate protects the moment of deploy. Drift detection protects the weeks in between. You need both, and they’re built from different machinery. Most teams “do evals” the way they once…
The trick to putting LLM agents in high-stakes systems isn’t a smarter model — it’s containing the model to one node so the rest of the system is ordinary, testable code. Here are the structural moves, with the contracts and types to implement them. Demos love autonomous agents that loop, call tools, and “figure it out.” Production hates them. The moment an agent’s behavior depends on which path the model wandered down today, you can’t test it, can’t audit it, and can’t let it touch anything that matters. In a regulated or high-consequence system — money movement, healthcare, infrastructure —…
Coding agents are great for an afternoon and mediocre for a quarter. Here’s the small set of files and rules — with the actual configs — that keeps quality from decaying across months and thousands of edits, whether you run one agent or a fleet of them on the same repo.
Tags: python architecture oop design-patterns distributed-systems When engineering distributed monitoring agents or designing low-latency health-checking pipelines, separating centralized governance from autonomous edge execution is essential. I designed the Prime-Sentinel Command (PSC) architecture as an object-oriented master-agent pattern to coordinate edge diagnostic nodes (Sentinels) via a centralized orchestrator (Prime). Below is an architectural walkthrough and minimal reference implementation for engineers looking to build similar decoupled telemetry collectors. Many diagnostic…
Ryan chats with Erin Yepis, Senior Analyst at Stack Overflow, about the results from this year’s Annual Developer Survey, including the overwhelming daily usage of AI coding assistants despite lingering developer trust issues, the critical role of well-organized documentation in providing verifiable context to mitigate AI hallucinations, and the evolving ways developers are shifting away from active community posting in favor of passive knowledge consumption.
Below, we’ll highlight some of the results we found interesting about what's going in the life of technologists, their technologies, AI usage, and more.
Ryan chats with Julien Verlaguet, CEO at Skip Labs, about finding the balance between human tolerance and tooling constraints, the spectrum of typed programming languages, and building cost-effective tooling for AI agents.
This analysis compares the 2024 and 2025 Developer Survey data across three connected stories: the evolution of AI, humans at work, and demographics and community.
AI can make the first part remarkably fast. It can find the page, the discussion and the person who might know. The harder work begins when those sources disagree, or when they become stale.
Stack Internal transforms your daily work into a living memory that’s shared with the rest of your team. Now anyone can create and share their knowledge in a Stack Internal workspace for free by visiting stackinternal.com.
We are on the precipice of brand new results from the 2026 Developer Survey. We have 15 years of data and insights from this survey at our fingertips that shows more than just what one year’s collection of insights will explore. A year-over-year look at the Developer Survey offers the long view of what is changing in the broader technology ecosystem. This analysis compares the 2024 and 2025 Developer Survey data across three connected stories: the evolution of AI, humans at work, and demographics and community. The survey results reveal what the headline news cannot: how people actually fold…
Ryan sits down with Div Garg, CEO at AGI Inc., to talk about running AI agents entirely on mobile devices, optimizing models for edge computing chips, and building safety mechanisms into autonomous app interactions.
In the rush to adopt AI and automation, many teams implement human-in-the-loop (HITL) frameworks. They believe that involving a person in the process solves the problems with reliability, quality, and trust. But as we’ve learned from real engineering workflows and integrations, the story isn’t that easy. In some contexts, humans-in-the-loop do improve outcomes, but in others, they can unintentionally become bottlenecks that limit speed, scalability, and innovation. In this post, we’ll analyze when human-in-the-loop is truly valuable, when it slows systems down, and how to strike the right…
Read at the source
Your visit, your choice.
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.