Perplexity has launched Photon, an in-house retrieval and ranking engine written in Rust, to replace the open-source fork it previously maintained. The new system now processes all production traffic and powers a new Fast Search mode for the Perplexity Search API. Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95.
In this article
Developers can access this as a hosted API by setting search_type: "fast" on the POST /search endpoint. The cost is $1 per 1,000 requests. The engine itself is not open source, so self-hosting is not an option.
Why Perplexity Replaced its Old Engine
The legacy engine hit three specific limits as the index grew larger.
- Tail latency: Production p99 sat near 800 ms. The dataset exceeded RAM, so
mlockwas not an option. Cold reads triggered major page faults that stalled queries. - Merge spikes: During disk index fusion, p99 climbed to about 1.2 s for 10 to 15 minutes.
- Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also raised the share of partial responses.
The Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.
How Photon Works
A load balancer routes each request to a Photon broker. The broker fans out to a shard group and watches for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and fetches key document fields.
- Adaptive posting lists: Short lists sit inline within a single page. Longer lists split into blocks of fixed document ID ranges. Sparse blocks store sorted offset arrays and use galloping search. Dense blocks use bitmaps, so membership becomes a single bit lookup.
- Budgeted traversal: A WAND-like algorithm splits lists into driving lists and probe lists. Cheap presence checks bound each candidate’s maximum score first. Exact term frequencies are read only when a candidate can clear the threshold.
- Docblob records: Each document gets a compact record of frequencies, field masks, and positions. Terms use Elias-Fano encoding, so ranking decodes only the matched terms. Ranking a candidate needs just 1 lookup per document.
- Batched async reads: Record offsets are known upfront, so disk reads go out in batches through
io_uring. The cache checks the whole batch first. Readers take no locks, and eviction uses CLOCK instead of a shared LRU list. - Separate build and serve: Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes. A controller rotates serving groups one at a time and warms caches with replayed search-log queries.
A full web index now builds in a single-digit number of hours.
Production Results
- p99 retrieval and ranking latency fell from about 800 ms to about 65 ms. This covers Photon’s stages only.
- Photon runs on about 20% fewer serving machines than the old content nodes.
- It stores about 2.5x as much data per document, which Perplexity used to improve ranking quality.
- Pinning the same dataset with
mlockwould need an estimated 4.6x the resident memory Photon uses today. - Index version switches no longer cause latency spikes.
Fast Search: Speed and Cost for Agents
Fast Search pairs Photon with lighter ranking tuned for agentic workflows. Perplexity tested it on six benchmarks: WideSearch, BrowseComp, DSQA, FRAMES, SEAL-0, and SEAL-Hard. Across 3,554 tasks, Fast scored 64.3% at $59.73 in estimated model-plus-search cost. The default preset scored 64.0% at $187.60, so Fast was about 68% cheaper.
The trade-off shows up in broader search quality. On internal long-tail benchmarks, relevance (DCG) fell from 2.45 to 2.21. Answer availability dropped from 0.596 to 0.567, a loss of 2.9 percentage points. Perplexity recommends Fast for day-to-day agent loops and the default for hard, ambiguous queries.
On Python SDK 0.43.4 and 0.43.5, pass
extra_body={"search_type": "fast"}per the docs.
Fast Search vs Closest Competitors
| Feature | Perplexity Fast Search | Exa Instant | Parallel Search Turbo | Tavily ultra-fast |
|---|---|---|---|---|
| Request parameter | search_type: "fast" | type: "instant" | mode: "turbo" | search_depth: "ultra-fast" |
| Vendor-reported latency | 160 ms p50, 230 ms p95 | ~250 ms typical; sub-200 ms at launch | ~200 ms | No figure published; lowest-latency depth |
| List price per 1K requests | $1 | $4 for up to 10 results | $1 | 1 credit: $8 pay-as-you-go, $5 to $7.50 on plans |
| Results per request | 1 to 20 | 10 in base price, $1 per 1K per extra result | Not specified | Not specified |
| Known limits | Lower relevance than default preset | Extra results billed separately | English and Japanese queries only | Lower relevance than other depths |
| Launched | Sep 24, 2026 | Feb 12, 2026 | Jul 13, 2026 | Jan 5, 2026 |
All latency figures are vendor-reported under different setups, so they are not like-for-like.
What it means
For developers building agents, this offers a specific trade-off. You can reduce costs by roughly two-thirds for standard tasks, but you must accept slightly lower relevance on complex, ambiguous queries. The default preset remains the safer choice for those difficult questions.




