In July, we opened the doors of our Düsseldorf campus for this year’s QA Meetup, bringing together around 90 people from testing, engineering, and beyond. Hosting this meetup has become a bit of a tradition for our QA team over the past few years, and it’s been a great success every time. This year, we had a great lineup of three speakers: Andrei Khabarov from Qase.io, Alexei Vinogradov from Curio IT, and our own Renjith Rajasekharan from the Backend QA team. Two of the talks looked at how AI is changing software testing. One through adoption metrics, the other through an AI agent built to…
Every year, our tech community gathers for two days to learn, share knowledge, and connect. Over 200 tech talents, all in the same room, presenting their work, celebrating wins, and learning from each other through talks that showcase what’s being done across trivago’s tech landscape. Over 200 tech talents. 88 presenters. 2 days. One theme running through all of it: Frictionless. Before tech, we focus on people This year’s event kicked off with our Chief Technology Officer Ioannis doing something unexpected: putting down the slides and asking everyone to stand up. ”Find a person you don’t…
This post walks through a layered performance investigation that cut PSE-kafka’s infrastructure costs by 83% and ended a run of 19 P1 incidents. Problem Background PSE-kafka (price-search-engine-kafka) is an internal microservice. It consumes hotel price Kafka messages and pushes ads to an external Ads service, which drives real revenue for trivago. But the service had a few persistent problems. CPU usage in production sat very low, around 10%. On startup, it took a long time to join its Kafka consumer group, which dragged out every deployment. It suffered from high consumer lag. The replica…
The problem Does 99.8% represent a good success rate in end-to-end test automation of a web application? What about 99.9%? Maybe yes, but when you have a test suite that has up to 1,000 scenarios, it means often having one or two failing tests, due to the well known issue of test flakiness. Maybe even more than that, when the environment has a bad day. Our preview and stage systems have high availability, but they are obviously not as reliable and performant as production systems, as for example many services might have some increased latency in their responses. So when you trigger your…
Most GraphQL Gateway discussions focus on public-facing APIs and multi-client architectures. This article explores a different axis—using GraphQL Mesh to stitch internal services into a unified gateway powering admin tooling. We share implementation details, honest challenges from six years in production, and a forward-looking perspective on how AI agents could leverage the unified graph.
The output of most automated accessibility tests is a long list of violations. This format, while comprehensive, makes it difficult to distinguish new issues introduced in a feature branch from long-standing technical debt. It doesn’t clearly show if accessibility is improving, and it doesn’t help prevent regressions. With regulations like the European Accessibility Act (EAA) making digital accessibility a legal requirement, teams need a more effective process than simply reviewing an ever-growing list. Most automated accessibility testing tools answer “What’s wrong right now?” but they fall…
Introduction / Context Kafka sits at the heart of how we move data between systems at trivago. Many teams publish changes to Kafka, and downstream services consume those changes to keep user-facing features up to date—things like accommodation reviews, highlights, and other derived attributes. To make that possible at low latency, we run a fleet of Kafka consumers we call sinks. Each sink takes events from Kafka, applies the necessary business logic (filtering, normalization, policy checks, etc.), and writes the result into a service-local database as a materialized view. This pattern works…
Migration projects can be hard, especially when we were not around when the original projects were built. We migrated our images infrastructure to Google Cloud which was spread across multiple environments and here is how we did that.
Learning about risk-taking, sentiment analysis, cybersecurity, and psychological safety in one evening? That’s quite a mix, isn’t it? That’s what the audience got offered at the Women in Tech meetup, which took place on February 29th in trivago’s office space. Together with iteratec, trivago had the pleasure of bringing together the vibrant community of female tech talents and their allies on the occasion of International Women’s Day, for an evening featuring stories about risk, safety, and continuous improvement. Aida Orujova and Sophia Breth guided the crowd through the evening in which 60…
Anomaly detection for time series is like finding unusual events in a sequence of data over time. It helps identify outliers or deviations from the expected pattern, signaling potential issues or anomalies in the dataset. This is the theory, but how does it translate into practical implementation for real business needs?
TL;DR During the development of customer-facing applications, time is crucial, especially when it comes to testing and analyzing changes before accepting them in production. This blog post explores how we developed a Java-based reactive tool to simulate production requests, that allows us to have quicker hints about the effects of changes introduced and be more confident about the hypotheses that are formulated. As a long term vision, we wish to significantly reduce the A/B testing time and ensure seamless transitions. Motivation At trivago, our goal is to provide our users with the best…
Retry on failure, good or bad? Why should you retry all tests on failure? Why not? This article will not go into details, listing pros and cons of each approach. There are already enough resources on the Web about the topic, listing valid points for both opposing views. As trivago Hotel Search frontend QA team over the last years we tried to stay away from a brute-force retry policy for failures and we rather tried to execute test retries only in selected cases. Recently, when we switched to a Continuous Deployment approach for our new frontend Web application (which empowers developers to…
Read at the source
Your visit, your choice.
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.