Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase. We do…

By Vane September 30, 2026 3 min read
Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

Perplexity has launched Photon, an in-house retrieval and ranking engine written in Rust, to replace the open-source fork it previously maintained. The new system now processes all production traffic and powers a new Fast Search mode for the Perplexity Search API. Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95.

Developers can access this as a hosted API by setting search_type: "fast" on the POST /search endpoint. The cost is $1 per 1,000 requests. The engine itself is not open source, so self-hosting is not an option.

Why Perplexity Replaced its Old Engine

The legacy engine hit three specific limits as the index grew larger.

  • Tail latency: Production p99 sat near 800 ms. The dataset exceeded RAM, so mlock was not an option. Cold reads triggered major page faults that stalled queries.
  • Merge spikes: During disk index fusion, p99 climbed to about 1.2 s for 10 to 15 minutes.
  • Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also raised the share of partial responses.

The Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.

How Photon Works

A load balancer routes each request to a Photon broker. The broker fans out to a shard group and watches for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and fetches key document fields.

  • Adaptive posting lists: Short lists sit inline within a single page. Longer lists split into blocks of fixed document ID ranges. Sparse blocks store sorted offset arrays and use galloping search. Dense blocks use bitmaps, so membership becomes a single bit lookup.
  • Budgeted traversal: A WAND-like algorithm splits lists into driving lists and probe lists. Cheap presence checks bound each candidate’s maximum score first. Exact term frequencies are read only when a candidate can clear the threshold.
  • Docblob records: Each document gets a compact record of frequencies, field masks, and positions. Terms use Elias-Fano encoding, so ranking decodes only the matched terms. Ranking a candidate needs just 1 lookup per document.
  • Batched async reads: Record offsets are known upfront, so disk reads go out in batches through io_uring. The cache checks the whole batch first. Readers take no locks, and eviction uses CLOCK instead of a shared LRU list.
  • Separate build and serve: Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes. A controller rotates serving groups one at a time and warms caches with replayed search-log queries.

A full web index now builds in a single-digit number of hours.

Production Results

  • p99 retrieval and ranking latency fell from about 800 ms to about 65 ms. This covers Photon’s stages only.
  • Photon runs on about 20% fewer serving machines than the old content nodes.
  • It stores about 2.5x as much data per document, which Perplexity used to improve ranking quality.
  • Pinning the same dataset with mlock would need an estimated 4.6x the resident memory Photon uses today.
  • Index version switches no longer cause latency spikes.

Fast Search: Speed and Cost for Agents

Fast Search pairs Photon with lighter ranking tuned for agentic workflows. Perplexity tested it on six benchmarks: WideSearch, BrowseComp, DSQA, FRAMES, SEAL-0, and SEAL-Hard. Across 3,554 tasks, Fast scored 64.3% at $59.73 in estimated model-plus-search cost. The default preset scored 64.0% at $187.60, so Fast was about 68% cheaper.

The trade-off shows up in broader search quality. On internal long-tail benchmarks, relevance (DCG) fell from 2.45 to 2.21. Answer availability dropped from 0.596 to 0.567, a loss of 2.9 percentage points. Perplexity recommends Fast for day-to-day agent loops and the default for hard, ambiguous queries.

On Python SDK 0.43.4 and 0.43.5, pass extra_body={"search_type": "fast"} per the docs.

Fast Search vs Closest Competitors

FeaturePerplexity Fast SearchExa InstantParallel Search TurboTavily ultra-fast
Request parametersearch_type: "fast"type: "instant"mode: "turbo"search_depth: "ultra-fast"
Vendor-reported latency160 ms p50, 230 ms p95~250 ms typical; sub-200 ms at launch~200 msNo figure published; lowest-latency depth
List price per 1K requests$1$4 for up to 10 results$11 credit: $8 pay-as-you-go, $5 to $7.50 on plans
Results per request1 to 2010 in base price, $1 per 1K per extra resultNot specifiedNot specified
Known limitsLower relevance than default presetExtra results billed separatelyEnglish and Japanese queries onlyLower relevance than other depths
LaunchedSep 24, 2026Feb 12, 2026Jul 13, 2026Jan 5, 2026

All latency figures are vendor-reported under different setups, so they are not like-for-like.

What it means

For developers building agents, this offers a specific trade-off. You can reduce costs by roughly two-thirds for standard tasks, but you must accept slightly lower relevance on complex, ambiguous queries. The default preset remains the safer choice for those difficult questions.

Scroll to Top