Uber has slashed end-to-end search latency by 50% for Uber Eats after overhauling major portions of its search pipeline, the company reported. The improvements touch retrieval systems, feature hydration, ranking algorithms, advertising delivery, presentation layers, and underlying infrastructure. An agentic coding workflow helped engineers spot, test, and confirm further gains throughout the process.

The redesign started with a shift in how Uber measures delay. Rather than tracking backend API response time, the team began monitoring Above-the-Fold completion, defined as the interval until the first screen of results appears with images loaded. Pagination combined with server-side caching shrank the initial response, while asynchronous rendering let result items be processed at the same time. These adjustments improved Above-the-Fold latency by more than 200 milliseconds. Uber trimmed retrieval workload after discovering that tens of thousands of candidates underwent hydration before ranking, only to discard many later. Cutting low-value retrieval strategies saved roughly 120 milliseconds, and product-level embeddings dropped data lookups by more than 100 times, shaving another 50 milliseconds. Separating ranking hydration from presentation data reduced latency by over 100 milliseconds, while dependency removal saved 35 milliseconds and request hedging contributed 40 milliseconds. The advertising pathway was rebuilt with column-oriented bid data, in-memory access, and reduced serialization, trimming about 130 milliseconds. Further infrastructure changes included parallel encoding, smaller embeddings, connection management upgrades, and Go data structure tweaks to lighten garbage collection overhead.

The work has sparked discussion among engineers reviewing the project publicly. Anubhooti Nagar framed the performance challenge as "less about doing things faster and more about doing less work and avoiding unnecessary waiting," and pointed to Uber's Measure, Identify, Fix, Validate loop as a framework for ongoing performance gains. Pratik Dhanave emphasized that the outcome resulted from incremental optimization rather than a single architectural leap, calling it "no single big idea behind it, but a long list of careful decisions across the full stack." Vidya Pandey distilled the changes into three principles: Do less work, start work earlier, and remove unnecessary dependencies. Pandey also linked Uber's planned microbatching approach to techniques deployed in AI systems to cut synchronization between processing stages.

The improvements rest on Uber's existing search platform, which relies on Apache Lucene, Spark-based indexing, Kafka-based streaming updates, and a distributed serving layer, according to the company. Uber is now testing end-to-end microbatching, product-based retrieval, Zero Pass Ranking, and HTTP multipart streaming. Early trials of product-based search have delivered more than a 50% reduction in p99 latency. The planned changes let processing stages overlap instead of waiting for entire preceding stages to finish, pushing latency gains further. For platforms where every millisecond shapes conversion and revenue, this blueprint suggests that marginal improvements compound into strategic advantage when executed systematically. The real test will be whether competitors can replicate this discipline without the engineering depth and tooling Uber has built over years.