Developer ToolsPerformance

How Uber Eats Halved Search Latency: Six Techniques to Steal

Uber Engineering published a detailed breakdown this week of how they cut Uber Eats search latency by 50 percent — across six distinct layers of their stack, with specific millisecond savings per change. None of these techniques are Uber-specific. If you run a search pipeline, an API with complex hydration, or a backend that aggregates data before responding, at least half of these apply to you directly.

Fix Your Metric Before You Touch Your Code

The most important change Uber made had nothing to do with code. They stopped measuring backend API response time and started measuring Above-the-Fold (ATF) completion — the time from query submission until the first screen of results renders with images.

That reframing alone surfaced a 200ms win: pagination with server-side caching plus asynchronous template rendering dropped ATF latency immediately, before a single pipeline component was touched. They were optimizing backend p50 while users were waiting for the page to paint. If you have never audited what your latency metric actually corresponds to in user experience, that is where to start.

Three Principles, Six Areas, Hundreds of Milliseconds

Every search latency optimization Uber made maps to one of three principles: do less work, start work earlier, or remove unnecessary dependencies. Here is the breakdown across their six optimization areas:

AreaTechniqueSaving
PresentationPagination + async ATF rendering200ms+
RetrievalRemove low-value lexical strategies~120ms
RetrievalProduct-level embeddings (100x fewer lookups)~50ms
HydrationParallel ranking + presentation phases100ms+
HydrationRequest hedging on presentation layer~40ms
AdsColumn-oriented bid data + in-memory access~130ms combined
InfrastructureParallel encoding/decoding50ms+
InfrastructureEmbedding precision reduction (46% smaller)Halved DB query latency
InfrastructureParallel service mesh connectionsUp to 53ms
InfrastructureStack-allocated value types in GoGC overhead reduction

The Three Techniques Most Worth Stealing

Audit Your Retrieval Before You Optimize Anything Else

Uber discovered they were hydrating tens of thousands of candidates before ranking — and discarding most of them. Broad-recall lexical retrieval strategies were adding ~120ms while contributing negligible incremental recall. Removing them was free latency.

Product-level embeddings then allowed Uber to group similar items and reduce data lookups by over 100 times, cutting another 50ms from feature fetching. The lesson: before profiling your ranking or hydration code, count how many candidates enter your pipeline at each stage and how many exit. The difference is waste you are paying for in latency.

Split Your Hydration Into Two Phases

Ranking hydration (the data you need to score an item) and presentation hydration (the data you need to display it) are different things. Running them sequentially means your ranking stage is blocked waiting for display data it does not need yet. Uber split these into parallel phases and saved 100ms+ on the critical path.

Request hedging extended this further: for presentation hydration calls that go to external services, sending a second request to a different node if the first is slow cut aggregate tail latency by ~40ms. The cost is a redundant request on some fraction of calls — usually a reasonable trade for p95/p99 improvements.

Reduce Embedding Precision Without Cutting Accuracy

Uber reduced floating-point precision on their embeddings to 5 decimal places and switched to variable-length integer encoding. Embedding size dropped 46 percent. Database query latency halved. Semantic similarity results did not change meaningfully — five decimal places is sufficient for cosine distance comparisons. If you are storing dense embeddings for retrieval or similarity, check your current precision. Full float64 for approximate nearest-neighbor search is almost always overkill.

They Used an AI Agent to Find the Rest

After exhausting the obvious wins, Uber deployed an AI coding agent with custom engineering tools: the ability to pull live production latency profiles, identify the top bottlenecks by span, and draft code fixes for the most promising ones. The agent ran autonomously, catching redundant work on the critical path and unnecessary config fetches in the query loop that manual profiling had missed.

This is worth noting separately. AI-assisted performance engineering — give an agent profiling access, a latency target, and permission to draft patches — is a pattern now in production at Uber. It is reproducible. The tools to do this exist. If you are running a performance sprint, this is how it will look in 2026.

What Uber Is Working On Next

Uber is working on end-to-end microbatching (targeting another 100ms+ reduction), product-based search (early tests show 50%+ p99 improvement), and HTTP multipart streaming for progressive result rendering. The direction is consistent: push parallelism further and reduce blocking dependencies.

The Uber Engineering writeup is unusually detailed for an engineering blog post — production latency numbers this specific are rare. InfoQ’s coverage summarizes the highlights if you want a faster read. For deeper search architecture context, ByteBytego’s breakdown of the Uber Eats system is worth bookmarking alongside this one.

What to Apply This Week

You do not need Uber Eats scale for any of this to work. Measure what users experience rather than what your backend reports. Audit your retrieval discard rate before touching ranking code. Split hydration phases when ranking and presentation data can be fetched independently. Reduce embedding precision to what your similarity metric actually requires. Hedge tail latency on slow downstream calls. These techniques apply at 10,000 requests per day as readily as at 10 million — the numbers just get smaller.

Start with the metric. Everything else follows from measuring the right thing.

ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *