How We Took Search Latency from 300ms to 150ms
When I started working on search at RedDoorz, search was working fine at normal traffic. Average latency was around 300ms, which is not a bad number by itself. The problem showed up when traffic went up. GC graph was very jagged, GC was running every few seconds, Elasticsearch queries for pricing and availability were heavy, and under load both the application and Elasticsearch started building their own queues.
On top of this we were adding features every few months, new discount types, recommendations, pricing rules, and each one was adding one more condition or one more query to the search path.
So the goal was not really “make 300ms into 150ms”. The goal was to reduce the work done per search so we had some headroom when traffic spikes. 150ms is what we ended up with after doing that. Let’s go through what we did.
Elasticsearch upgrade and query cleanup
We were on Elasticsearch 5, so first thing was to upgrade to Elasticsearch 6.
After that we went through the search flow and the query structure. There was no single big problem here, just many small things which added up.
Redundant conditions. We had an availability flag and also an open_rooms field. open_rooms > 0 already tells you the room is available, so we dropped the flag and the extra condition.
Fields that are always searched together. Room type and date were separate fields, but search almost always needs the combination. So we indexed a combined value and searched on that directly instead of making ES evaluate two fields and combine them.
Rule we followed here: if something can be represented in the index, don’t make Elasticsearch compute it at query time.
Splitting pricing out of the property index
This was the bigger one. Pricing and availability were stored as nested documents inside the property index, and every search had to run nested queries to find properties matching the dates, room type, price and availability. As data and query complexity grew these nested queries became the most expensive part of the search.
We moved pricing and availability into a separate index. Search flow became:
- Query the pricing/availability index with the requested conditions
- Get the list of matching property ids
- Fetch property data for those ids
Yes, this is two queries instead of one. We expected it to be slower. It was faster, and more importantly it stayed stable under load. One expensive nested query turned out to be much worse than two simple ones.
Special query for single day searches
The availability check had to look at every room type for every day of the stay, so the query grew with room types × dates. When we looked at actual traffic, most of the searches were for a single day. So we wrote a separate, smaller query for single day searches and dropped all the conditions which only matter for multi-day stays. Multi-day searches still use the general query.
Keep the index small
Pricing and inventory for past dates was sitting in the active index for no reason. We added a daily job to delete expired inventory and pricing. Less data for ES to maintain and search over.
Shards
Our data was not big, but we were running on the default shard count. Each shard has its own cost in query coordination, memory, segments and merges. For our dataset we moved to a single primary shard where it made sense. Small change, but it removed overhead we were paying for nothing.
Multi search
The application was making several independent requests to ES for one user search. We combined them with _msearch wherever possible:
Before: After:
App → ES (search A) App → ES ─┬─ search A
App → ES (search B) ├─ search B
App → ES (search C) └─ search C
The searches still execute, but we saved the network round trips and request overhead.
Request cache and refresh interval
Hotel search traffic is not evenly spread. On popular dates or during promotions a lot of users run the same search. We enabled request caching so the same query is not recomputed when the underlying data hasn’t changed.
Caching raised the question of how fresh the data needs to be, which brought us to refresh intervals. Pricing and inventory change often, property data doesn’t. So we set them differently:
- pricing/inventory index:
refresh_interval: 5s - property index:
refresh_interval: 30s
Every refresh creates segments and background work, so there is no point refreshing property metadata every second.
Queue updates, then bulk write
We were already using bulk API for prices, reviews and inventory, but every event was still being sent to ES immediately. We put a queue in between:
Application events → Queue → collect → bulk update every ~30s → ES
Instead of a constant stream of small writes competing with searches, we now do one bulk write every 30 seconds. We traded a few seconds of indexing delay for a much cheaper write path.
Aggregations
Aggregations don’t need the full documents. We separated aggregation requests from the document fetch so we are not pulling property data when the request only needs counts. It also made aggregation results easier to cache.
Three APIs into one
The listing page shows three groups: properties the user booked before, recommended properties, and the rest. These were three separate API calls, and each one was making its own ES calls internally. So one user search was fanning out into a lot of requests.
We replaced this with a single search using function_score. Previously booked properties get a boost, recommended properties get a smaller boost, everything else falls back to normal ranking. One search, one response, same ordering as before, and a bunch of API and ES calls gone.
Pricing calculation in the application
By this point ES was not the only bottleneck. Pricing in the application was looping over the same data multiple times, once for promotions, once for discounts, once for pricing rules, once for special offers. We combined these so the data is iterated once and everything is calculated in that single pass. This cut CPU and helped a lot with the GC pressure.
Redis
Pricing also depends on data in Redis. We were holding large in-process caches in the application, which was part of the memory and GC problem. We dropped most of that and relied on Redis directly, using MGET and hashes to fetch everything needed in one or two calls instead of many individual GETs. Smaller process memory, less GC.
Recommendations: batch where possible, real-time where needed
Recommendation scores were being calculated during every search. Most of them didn’t need to be. We moved score generation to a batch job which writes scores ahead of time, and search just reads them.
Some recommendations did need real-time user behaviour, so those stayed in the search path. That’s a balance between freshness, personalization and latency, and we had to make that call per feature rather than once.
It never stops
This is the part worth saying honestly. We did not optimize once and move on. Every few months something new landed, new discount types, batch recommendation scores, then real-time recommendations, new ranking rules, new pricing rules. Each one was useful for the customer and each one added work to the search path.
So for every new feature the questions were the same: where does this computation happen, does it need to be synchronous, can it be cached, can it be precomputed, can ES do it cheaper, can it move out of the critical path. The job was never to say no to the feature. It was to decide where the expensive part runs.
Why bother if 300ms was fine
At normal traffic, 300ms is fine. But traffic is not normal during holidays, promotions and campaigns. When requests come in at several times the usual rate, they queue at the app layer, ES queries queue, CPU goes up, GC runs more often, and the expensive queries get slower because everything is contending. That 300ms search is not 300ms anymore.
A search that takes 150ms holds resources for half the time on the critical path. That is the headroom. The number going down was the side effect; the capacity to absorb spikes was the actual result.
Summary
Nothing on this list was magic on its own:
- ES 5 → 6
- Removed redundant conditions, merged always-together fields
- Moved pricing/availability out of the nested property index
- Separate query for single day searches
- Daily cleanup of expired inventory
- Single primary shard
_msearch, request cache, per-index refresh intervals- Queued updates with bulk writes every 30s
- Aggregations separated from document fetch
- Three listing APIs → one
function_scorequery - Single pass pricing calculation
- Redis
MGET/hashes instead of in-process caches - Precomputed recommendation scores
Together they got us from roughly 300ms to 150ms, and more importantly gave us room for the traffic spikes that actually matter.

