Speeding Up and Amplifying Experimentation Learnings Beyond the Platform
๐คAI Transparency Statement: This part of the article is 100% authored by Paul with only basic spelling and grammar checks.
Introduction
This was a project I ran in Fall and Winter of 2025 when working on Expedia Group’s internal online experimentation data platform. Although I am exiting the online experimentation domain, I think its a great example of showing how I worked not just to shape the product and features, but overcoming a tough operational challenge and narrative within how the organization itself operated.
Need more context? A/B experimentation has become the gold standard for objectively determining the effectiveness of an idea or feature. It puts subjective gut feel and intuition to the test by stating the goal, randomly offering the user competing experiences, and objectively measuring the results. Some organizations, like Expedia, opt to build their own platform to collect and analyze experimentation data, while others use a third-party tool like Kameleoon, Statsig or Datadog. Providing a tool or platform isn’t enough though, the people in the organization have to understand the craft as well.
I was approached by Katie Green at Kameleoon in Spring 2026 with an opportunity to be on their Unite Voices podcast, gave a synopsis of this project, and was told it was a good fit. I wrote up the experience (after my employment concluded), had AI assist in generating my talking points from it, and created below as well.
The podcast was published on June 17, 2026, and you can find it on Kameloon’s website. The core content of below overlaps heavily with the podcast, and reflects a full version of what I then broke down to circulate internally in our monthly newsletter, announcement channels, training and more.
๐คAI Transparency Statement: The rest of this article was co-authored with Claude Sonnet 4.6. I wrote the original long-form content, then had Claude propose a draft, which I then lightly edited.
Background and Problem
Across the organization, there was a persistent narrative: experimentation was moving too slow. The flywheel wasn’t spinning fast enough. Experimentation had become an anchor; a gate teams had to pass through. Teams poured months of hard work into a feature, only to reach the finish line and wait on an experiment to run, frequently walking away with inconclusive, conflicting, or otherwise unclear results. It was something you had to do, but real learnings rarely came out of it.
The natural reaction was to blame the platform. Features were missing. The data wasn’t right. There weren’t enough insights, not enough automated guidance. Because of this and because gut instinct tends to fill the void when data disappoints, proposals began circulating to bypass experimentation altogether, at least until the platform was “complete” or could provide a “100% automated readout.”
The problem with that framing is that neither of those conditions is ever truly achievable. Agreeing on what “complete” means in a large, evolving organization is nearly impossible, and any definition of “100% automated” would be obsolete the moment the business, its technology, or its users changed. These were rationalizations, not solutions.
The reality of running a decentralized experimentation program at scale is that the platform is only one piece of the puzzle. A healthy program depends on four interconnected elements, which I labeled the 4 P’s:
- People โ The practitioners who run experiments. They need to understand the platform, the process, and the purpose.
- Platform โ The system that collects data, surfaces results, and delivers the experimentation experience.
- Process โ The operational workflows that govern how experimentation fits into product development.
- Policy โ The documented standards and rules that define what good looks like.
Our platform had improved and matured significantly. We were in a polishing phase, filling in increasingly smaller gaps with a clear roadmap ahead. While support for niche and cutting-edge use cases still lagged, the reality for the majority of users, I estimated to be somewhere in the range of 70โ80%, was that the platform was ready. They should have been able to move faster.
That meant it was time to focus on process and people. Together, I often referred to this as the culture of experimentation, or more broadly, the operational habits baked into how products were built and shipped.
The challenge with decentralized programs in large, diverse organizations is that things drift. Without a strong center consistently setting standards, delivering them, and monitoring their health over time, organic variation takes over. That was our reality. The experimentation culture, and how it connected to the broader product operating model had drifted. Teams had lost clarity on when to start an experiment and how to scope it; how to write a meaningful hypothesis; what experimentation was actually for; and how to interpret results in a way that drove confident decisions.
What We Were Seeing
I spent a lot of time close to the experiments being designed and run. I operated in our support channels, exchanged DMs with experiment owners, reviewed experiments constantly while validating our own platform releases, and dropped into peer review discussions in meetings and Slack. The result was a large body of anecdotal signals gathered from hundreds of experiments over time, supplemented by internal behavioral data and usage analytics.
What I consistently observed fell into a few clear patterns:
Experiments were designed too large. Rather than validating something small and iterating quickly, teams routinely bundled 5, 10, or more major changes into a single test. This made it nearly impossible to instrument the right metrics โ and even harder to interpret results, since there was no clean way to attribute which change drove which outcome. Analysts were frequently pulled in for deep-dives just to separate signal from noise, and even then, the picture was often muddy.
Experiments launched too late. Related to the above, teams typically waited until a feature was fully built before testing it. By that point, the sunk cost was enormous. When results came back unfavorable, teams would bend over backwards to “save” an experiment rather than accept the outcome, because abandoning it meant admitting that months of work hadn’t paid off. The incentive structure worked against honest interpretation.
Hypotheses were vague. It was often unclear exactly what would change, why it would change, and what a successful outcome would look like. Without a well-formed hypothesis, it’s nearly impossible to know whether you achieved your goal.
Too many metrics were being used. Secondary metrics had proliferated. The more metrics included in a test, the higher the probability of false positives and contradictory results. Experiment readouts were frequently lit up like a Christmas tree as a mix of green and red across dozens of metrics, with no clear signal. Our platform supported one primary metric (which, per policy, should alone dictate rollout or rollback decisions), up to five secondary metrics to inform decisions and guide the next hypothesis, and a set of informational metrics to explain user behavior. In practice, those boundaries were routinely ignored.
Primary and secondary metrics were too insensitive โ and often irrelevant. Many metrics were selected because they mapped to financial reporting targets or org-level goals, not because they were actually connected to what the experiment was testing. Running a test on a specific feature interaction and measuring success in downstream booking conversions, a user journey that could span weeks for a travel booking site meant most experiments were structurally destined to be inconclusive, or required impractically large sample sizes just to reach statistical significance.
Negative outcomes were underrepresented. Experimenters tended to design for success. Edge cases like increased error rates, longer load times, or unexpected regressions were rarely instrumented upfront. While our circuit breaker and automated guardrail metrics covered some of these scenarios, the teams closest to the product often knew the most sensitive indicators of something going wrong, but those rarely made it into experiment designs.
The through line across all of these observations was this: experimentation for learning had been largely abandoned. It had become exclusively a validation exercise. It was a mechanism for confirming decisions already made, mapped to financial and organizational goals rather than to the actual expected behavior of users interacting with a new feature. This made experiments run longer, produced inconclusive results, and encouraged teams to build features fully before testing them, rather than using experimentation to shape what got built.
Hypothesis
If the problems were process and behavior, the solutions needed to be too. The core thesis was straightforward:
- Test sooner โ earlier in the development cycle, before full builds are committed.
- Test smaller โ fewer changes per experiment, sharper scope.
- Use more sensitive metrics โ ones that are proximate to the actual change being made, rather than far-downstream financial indicators. Guardrail metrics and circuit breakers exist to protect against harm to those downstream goals; primary metrics should be tuned to detect what the experiment is actually measuring.
- Use fewer decision metrics โ a small, well-chosen set beats a sprawling dashboard.
- Design for negative outcomes โ at minimum, informational metrics that detect things going wrong should be present from the start.
What We Did
Building the Case
I started by writing up a structured version of the problem and the proposed approach, capturing the observations above, pulling in supporting references, and sketching out illustrative examples. I used that document to collect peer feedback, refine the argument, and begin socializing it.
I knew most teams wouldn’t be immediately open to changing their approach. They were heads-down against org goals and couldn’t afford what felt like a process experiment on top of their product work. Finding the right first partner mattered.
Finding the Right Pilot Team
We identified a team that was actively feeling the pain. They’d done the work to validate that their focus areas mattered to users and were directionally supported by industry signals, but struggled to prove it through how tests were currently being structured. They were hungry for a better way. It didn’t take more than a brief conversation to get alignment. A PM and their analyst were nominated and immediately bought in.
Designing the Pilot
We focused the pilot on two parallel workstreams for this team. The PM shared planning documents and PRDs so I could understand not just what was being built, but why. Both were essential to designing experiments that could actually answer meaningful questions. I also reviewed a set of their recent and in-progress experiments to calibrate where they were starting from.
We ultimately focused on two streams: one aimed at improving a specific feature’s performance (notoriously hard to tie to financial goals, even when it clearly mattered to users), and another involving a bundle of feature improvements that the team would have traditionally waited to fully complete before testing as a single experiment.
Running It
I wrote an initial general strategy guide and then translated it into detailed experiment plans for each workstream. My first draft took about an hour. Breaking out of deeply ingrained habits is harder than it sounds. Whittling down from the 11โ13 metrics typically used to a focused set of 3โ5 required deliberate effort and felt uncomfortable. It was the right kind of uncomfortable.
Alongside this, we identified that some of the metrics we needed didn’t exist yet. Creating new, purpose-built metrics for this team’s specific use case forced us to work through the metric onboarding process. This also turned out to be a useful exercise in building empathy for a distinct category of platform pain points that later fed into our roadmap.
I reviewed the initial proposal with the PM and analyst in a full working session. There was healthy debate. They learned about platform capabilities that they hadn’t fully understood before to better attribute the experiment to the right moment in the user journey, and automated guardrail protections.
The second experiment plan came together in about 30 minutes. The muscles were already developing, and times dropped from there.
Learning in the Field
The first test launched and failed within days โ not due to UX issues or bugs, but due to instrumentation problems with the new metrics setup. That failure surfaced early what would have eventually surfaced later and caused greater delays. It was valuable information, quickly.
The test relaunched. We used some proxy metrics and supplemented with non-automated data while the new instrumentation continued to mature. A few debugging metrics were added to distinguish experiment-level issues from actual user-facing problems. It ran to completion and resulted in a temporary rollback on the primary goal. But the learnings were rich. Metrics expected to be highly correlated with each other showed no correlation at all. Early data on actual user behavior patterns started to emerge, including some counterintuitive findings about what outcomes even constituted success.
A third iteration spun up with lower overhead and less direct involvement from me. It ran again and this time met its primary goal. A rollout. And beyond the rollout, the team had accumulated a clear picture of which metrics were worth keeping, which could be retired, and where new goals could be built to amplify what they were learning.
The second workstream ran in parallel and followed a similar arc, with its own version of wins that the team was quick to share out and celebrate.
The Differences Were Clear
By the end of the pilot, the contrast with the previous approach was concrete:
- Test cycles dropped from 4โ6 weeks each to 2 weeks, the minimum specified in our playbook. There was early evidence that even that threshold might not be universally necessary depending on the test type.
- Analysis time dropped from days to hours. No analyst interventions required. Results were readable without a deep-dive.
- Beyond validating features, the team was generating genuine learnings about how users behave that accelerated ideation on future features and informed better experiment designs going forward.
Perhaps most importantly, everyone involved was optimistic. They had acutely felt the pain of the status quo. They just hadn’t known exactly how to address it, or felt sufficiently empowered to swim against the current on their own. The pilot gave them a foundation to build from.
Scaling the Approach
There was real excitement around what the pilot had demonstrated. Small proof-of-concept, yes, but proven once. The goal became creating momentum.
We used our internal experimentation newsletter to share the story from the PM’s point of view. We opened a dedicated Slack channel for the initiative, giving teams a place to ask questions and build confidence. We offered open office hours for a few weeks. And I wrote up the work in a case study format, then used that to develop a detailed guide on how to design experiments this way.
From there, I partnered with a completely different team. I adapted the content into a workshop format: slides, discussion, and guided exercises. I worked with their VP and a senior PM advocate to carve out time, frame the importance, and sharpen the content. We delivered a full afternoon session to PMs, analysts, and designers. The team was fully engaged. They understood the concepts and wanted to apply them. The barriers weren’t philosophical, they were practical: comfort with the platform, and confidence in how to make the pivot themselves.
I was in the process of working with internal learning and development teams to build this into a scalable, broadly available curriculum, something that wouldn’t require me to personally deliver and monitor every session when my time with Expedia Group ended. I can’t speak to the full downstream impact, but the seeds were planted, and the appetite was real.
What This Exposed
The most important finding wasn’t in the data. It was in the dynamic.
There was nothing uniquely broken about the tools or the people. There was no fundamental reason experiments had to run the way they had been running. The practices had drifted organically, reinforced by habit and the path of least resistance. What teams needed wasn’t a better platform, it was a clearer framework, permission to do things differently, and a tangible proof point that a better way was possible.
The appetite for change was there. It just needed a spark.
For us as an experimentation platform team, the learnings were just as valuable on the operational side. Every conversation improved the documentation and guidance. Working through the metric creation process with teams revealed new friction points that hadn’t been visible before and made a more compelling case for a category of work we’d been discussing but hadn’t fully prioritized: developing more sensitive proxy metrics that correlated strongly with financial and organizational goals without requiring teams to measure those goals directly in every experiment. It also helped us understand how teams were debugging issues outside the platform and workarounds we could bring in-platform to give experiment owners faster, more confident paths to decisions.
The human element needed an update and upgrade. Once we addressed that, everything else was unlocked to move faster.