AI is making web scraping faster - but dependable data takes more than working code. Quality, provenance, self-repair, access and responsibility are the new edge.
The famously awkward security mechanism has quietly become an invisible behavioral scoring system. Two vendors enable 89% of CAPTCHAs, and the smallest sites use it most.
Recordings, slides, and a rundown of Zyte's second Developer Community Meetup with Humanbound: a live prompt injection attack on a price agent, Zyte CDP managed browsers, and the launch of harness-run.
Scrapy 2.19's RemoteControl extension runs a small authenticated HTTP server inside every crawl. How it starts, how it's secured, what the code you send can reach, and what you can do with it on a live crawl, plus the traps and what changes under scrapy-zyte-api.
Scrapy MCP connects an AI agent to a running Scrapy crawl, using the new Remote Control extension in Scrapy 2.19 to let the agent discover jobs, check status and run Python against the live crawler, all without restarting it.
One Scrapy spider, three browser setups: stock Playwright, Patchright via the new PLAYWRIGHT_BROWSER_PROVIDER hook, and a remote browser on Zyte over CDP. Same spider and selectors throughout, only the browser changes.
A Scrapy pipeline that asks a fast, calibrated AI model whether each scraped field still looks real, and stops the crawl when too many don't. What it caught, what it misses, and whether building it was worth it.
Jev by TypeFace AI cannot generate a string, so it cannot extract a field. What it is, how it differs from an LLM, and the one job it earns in a scraping pipeline.
Only 18.5% of top sites run dedicated antibot, but every one of them chose to. Why it's the most intentional barrier in the stack, and the strongest signal of a hardened site.
uv 0.12 changed several defaults that Python developers rely on—from how `uv init` structures projects to how hash checking and lockfiles behave. This practical cheatsheet tests the differences across versions and highlights the command-line traps most likely to surprise you.
What if your AI agent could use a real browser to search, click, compare products, and collect data, without you writing a selector for every step? By connecting it to a remote Zyte CDP browser, you can turn plain-language instructions into practical browser automation.
It was just a silly game side project. But, as I grew my sim to 3,000 matches, I levelled up on edge routing, flattening memory spikes and hardcore monitoring.
Real web is becoming hostile for AI Agents. The page your agent scrapes now could be a potential attack surface. Read more and join Zyte's virtual meet-up to see it in action.
Running Playwright at scale: connecting to the Zyte CDP browserBlogRunning Playwright at scale: connecting to the Zyte CDP browserArticleTutorial / How-to Running Playwright at scale: connecting to the Zyte CDP browser John Rooney · Developer Engagement Manager September 7, 2026 Try Zyte API Build your first scraper in minutes Free trial, no credit card. From a single request to production in an afternoon. Get started John Rooney Developer Engagement Manager John is the Developer Engagement Manager at Zyte, working closely with the community, creating content and helping developers learn web…
Nearly four in 10 of the world's top sites are now closed to well-behaved AI crawlers. Here's what operators actually do with robots.txt, and which AI agents they block or welcome.
marimo is a reactive Python notebook that reruns only affected cells. Learn to scrape web data with Zyte API, chart prices, and cache costly API calls.
Three in four of the world's top sites publish a robots.txt, yet few name individual crawlers. A look inside the web's advisory layer: coverage, sophistication and limits.
Read at the source
Your visit, your choice.
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.