pi vs pi-fff: measuring what the extension actually costs
I have been running @ff-labs/pi-fff in pi for a while, mostly on the strength of the claim that fuzzy search over a pre-built in-process index beats spawning rg and fd on every tool call. That claim is easy to believe and easy to test, so I tested it.
The setup: pi 0.99.2, @ff-labs/pi-fff 0.11.0, CachyOS on an Intel Core Ultra X7 358H with an NVMe drive. Fresh process per run, SDK mode with no LLM calls, PI_OFFLINE=1. Numbers are medians across runs. The two conditions differ only by the packages: ["npm:@ff-labs/pi-fff"] entry in ~/.pi/agent/settings.json. For the "without" side I stripped that line in memory for the SDK tests and rewrote then restored the file for the CLI test. Everything else, other extensions, settings, model config, stayed identical.
Short version: with pi-fff, startup is about 40 ms slower (paired median), session start costs 17 to 51 MB more RSS, and every model request carries 1,031 extra tokens. In exchange, grep runs 6 to 33 times faster, find 12 to 16 times faster, and grep at $HOME 128 times faster. Those ratios sound dramatic, but the absolute savings are a few milliseconds per call. The cost that actually matters is the token count, because it repeats on every request for the whole session.
CLI startup
pi --mode rpc spawn to ready:
| condition | runs | min | median | mean | max |
|---|---|---|---|---|---|
| with | 25 | 217.8 | 266.4 | 258.2 | 291.8 |
| without | 25 | 190.7 | 216.9 | 217.4 | 243.7 |
Taken at face value, pi-fff adds 49.5 ms (22.8%) to CLI startup. I also ran the conditions interleaved, one with and one without per round, so drift cancels out when the runs are paired:
paired with − without |
n | median | Q1 | Q3 | rounds where with > without |
|---|---|---|---|---|---|
| 25 | 39.7 ms | 25.4 ms | 63.2 ms | 23/25 |
The paired median of 39.7 ms is the number I trust, and with was slower in 23 of the 25 rounds.
Session start, memory, and per-request payload
A few phase names first. reload is resource and package discovery, which is where pi-fff gets resolved and loaded. create is session construction. bind is session_start, which for pi-fff opens and builds the FFF file index for the current directory. rss is process memory right after startup, since the index lives in-process. payload is the system prompt plus the active tool schemas sent on every model request (approx tokens = chars / 4).
| cwd | condition | runs | reload | create | startup | session_start | rss after start | rss after tool runs | payload chars | ≈tokens |
|---|---|---|---|---|---|---|---|---|---|---|
| pi tree (4.7k files) | with | 5 | 88.5 | 16.7 | 105.1 | 62.5 | 151 MB | 230 MB | 19302 | 4825 |
| pi tree (4.7k files) | without | 5 | 78.9 | 16.1 | 96.3 | 0.7 | 134 MB | 168 MB | 15178 | 3794 |
| phoronix (11k files) | with | 5 | 84.4 | 15.8 | 101.6 | 62.0 | 155 MB | 195 MB | 19285 | 4821 |
| phoronix (11k files) | without | 5 | 79.0 | 16.4 | 95.8 | 0.7 | 134 MB | 145 MB | 15161 | 3790 |
| $HOME (1.2M files) | with | 5 | 95.9 | 17.2 | 111.8 | 67.0 | 186 MB | 376 MB | 19253 | 4813 |
| $HOME (1.2M files) | without | 5 | 92.6 | 17.1 | 109.5 | 0.7 | 134 MB | 146 MB | 15129 | 3782 |
| cwd | extra session_start time |
extra RSS |
|---|---|---|
| pi tree (4.7k files) | +61.8 ms | +16.8 MB |
| phoronix (11k files) | +61.3 ms | +21.2 MB |
| $HOME (1.2M files) | +66.3 ms | +51.3 MB |
The interesting line is the payload. pi-fff adds 4,124 characters, roughly 1,031 tokens, to every model request: 15,129 chars becomes 19,253 at $HOME. Two extra tool schemas, plus their guidelines in the system prompt. You pay that on the first request, the tenth, and the last one before you close the terminal.
Tool-call latency
Same session, same query, same result set. grep and find are pi's built-ins and spawn rg/fd per call; ffgrep and fffind are the pi-fff replacements that work over the pre-built index. cold is the first call of the run, median covers the rest. Only with runs are counted, because the fff tools do not exist otherwise. out is the character count of the returned result, included so you can see that both sides did comparable work.
| cwd | query | tool | cold ms | warm median ms | result chars | errors |
|---|---|---|---|---|---|---|
| pi tree (4.7k files) | find-fuzzy | <builtin> |
n/a | n/a | n/a | 5 skipped |
| pi tree (4.7k files) | find-fuzzy | fffind |
1.5 | 1.3 | 1827.0 | 0/5 |
| pi tree (4.7k files) | find-glob-all | fffind |
2.1 | 0.8 | 34577.0 | 0/5 |
| pi tree (4.7k files) | find-glob-all | find |
12.3 | 9.9 | 51279.0 | 0/5 |
| pi tree (4.7k files) | grep-literal-common | ffgrep |
2.3 | 0.2 | 3301.0 | 0/5 |
| pi tree (4.7k files) | grep-literal-common | grep |
8.6 | 4.7 | 15718.0 | 0/5 |
| pi tree (4.7k files) | grep-literal-identifier | ffgrep |
0.5 | 0.2 | 2903.0 | 0/5 |
| pi tree (4.7k files) | grep-literal-identifier | grep |
6.4 | 5.9 | 15808.0 | 0/5 |
| pi tree (4.7k files) | grep-literal-rare | ffgrep |
5.6 | 4.5 | 16.0 | 0/5 |
| pi tree (4.7k files) | grep-literal-rare | grep |
8.5 | 8.8 | 16.0 | 0/5 |
| pi tree (4.7k files) | grep-regex | ffgrep |
0.7 | 0.3 | 5346.0 | 0/5 |
| pi tree (4.7k files) | grep-regex | grep |
8.0 | 6.8 | 27012.0 | 0/5 |
| pi tree (4.7k files) | grep-with-path-filter | ffgrep |
0.7 | 0.3 | 1036.0 | 0/5 |
| pi tree (4.7k files) | grep-with-path-filter | grep |
4.1 | 3.7 | 13569.0 | 0/5 |
| phoronix (11k files) | find-fuzzy | <builtin> |
n/a | n/a | n/a | 5 skipped |
| phoronix (11k files) | find-fuzzy | fffind |
0.8 | 0.6 | 1897.0 | 0/5 |
| phoronix (11k files) | find-glob-all | fffind |
3.1 | 0.8 | 18781.0 | 0/5 |
| phoronix (11k files) | find-glob-all | find |
13.1 | 12.4 | 18781.0 | 0/5 |
| phoronix (11k files) | grep-literal-common | ffgrep |
2.3 | 0.2 | 1696.0 | 0/5 |
| phoronix (11k files) | grep-literal-common | grep |
9.2 | 3.6 | 14131.0 | 0/5 |
| phoronix (11k files) | grep-literal-identifier | ffgrep |
0.6 | 0.5 | 3437.0 | 0/5 |
| phoronix (11k files) | grep-literal-identifier | grep |
3.7 | 3.1 | 12261.0 | 0/5 |
| phoronix (11k files) | grep-literal-rare | ffgrep |
5.9 | 4.6 | 16.0 | 0/5 |
| phoronix (11k files) | grep-literal-rare | grep |
13.1 | 12.4 | 16.0 | 0/5 |
| phoronix (11k files) | grep-regex | ffgrep |
0.8 | 0.4 | 1410.0 | 0/5 |
| phoronix (11k files) | grep-regex | grep |
13.0 | 13.0 | 1582.0 | 0/5 |
| phoronix (11k files) | grep-with-path-filter | ffgrep |
1.3 | 0.5 | 1843.0 | 0/5 |
| phoronix (11k files) | grep-with-path-filter | grep |
3.4 | 3.2 | 10698.0 | 0/5 |
| $HOME (1.2M files) | find-glob-all | fffind |
6.3 | 1.7 | 55638.0 | 0/5 |
| $HOME (1.2M files) | find-glob-all | find |
24.5 | 19.9 | 51253.0 | 0/5 |
| $HOME (1.2M files) | grep-literal-identifier | ffgrep |
59.1 | 4.1 | 2120.0 | 0/5 |
| $HOME (1.2M files) | grep-literal-identifier | grep |
604.4 | 526.4 | 39638.0 | 0/5 |
Built-in versus fff, warm medians
| cwd | query | built-in ms | fff ms | speedup |
|---|---|---|---|---|
| pi tree (4.7k files) | find-glob-all | 9.9 | 0.8 | 12.4x |
| pi tree (4.7k files) | grep-literal-common | 4.7 | 0.2 | 23.5x |
| pi tree (4.7k files) | grep-literal-identifier | 5.9 | 0.2 | 29.5x |
| pi tree (4.7k files) | grep-literal-rare | 8.8 | 4.5 | 2.0x |
| pi tree (4.7k files) | grep-regex | 6.8 | 0.3 | 22.7x |
| pi tree (4.7k files) | grep-with-path-filter | 3.7 | 0.3 | 12.3x |
| phoronix (11k files) | find-glob-all | 12.4 | 0.8 | 15.5x |
| phoronix (11k files) | grep-literal-common | 3.6 | 0.2 | 18.0x |
| phoronix (11k files) | grep-literal-identifier | 3.1 | 0.5 | 6.2x |
| phoronix (11k files) | grep-literal-rare | 12.4 | 4.6 | 2.7x |
| phoronix (11k files) | grep-regex | 13.0 | 0.4 | 32.5x |
| phoronix (11k files) | grep-with-path-filter | 3.2 | 0.5 | 6.4x |
| $HOME (1.2M files) | find-glob-all | 19.9 | 1.7 | 11.7x |
| $HOME (1.2M files) | grep-literal-identifier | 526.4 | 4.1 | 128.4x |
Things worth knowing before you turn it on
The index covers non-hidden files, the same population as rg --files: 24,266 entries at $HOME out of 1,212,615 files on disk there, 11,032 at the phoronix tree (11,034 on disk), and 2,973 in the pi tree (4,709 on disk, with hidden and ignored files skipped). The index is built during session_start and is complete immediately. I probed at +0 s, +3 s, +10 s and +30 s and got the same 24,266 entries every time; a full listing takes 22 to 32 ms.
Memory grows with use. At $HOME RSS went from 196 MB to 240 MB right after the first full listing, then settled near 200 MB. A longer session with about 25 searches and globs finished at 429 MB, which is FFF's in-process content cache. The built-in tools hand that data to rg and fd child processes instead, so the pi process itself stays in the 135 to 200 MB range.
Cold and warm differ a lot on the fff side. The first ffgrep after session start costs about 59 ms at $HOME while the content cache is empty, then drops to around 4 ms. rg never gets faster: 526 ms warm.
A breakdown of where the CLI time goes: about 10 ms is package resolution and resource discovery, which shows up as the reload delta. The remaining 30 ms or so fits with loading and initializing the extension module itself, a 54 KB TypeScript entry plus its native libfff_c.so FFI binding.
One more caveat on scope. There are no LLM calls anywhere in these runs, so end-to-end turn latency, where tool speed usually gets dwarfed by model latency, is out of scope. This measures pi's own overhead.
Method notes
- Every paired query goes to both implementations with equivalent parameters, and
find-glob-allreturns byte-identical output fromfdandfffindon the phoronix tree. grepandfindresults are not byte-identical toffgrepandfffindbecause of different caps, grouping and ranking, but both sides scan the same tree. See theresult charscolumn.find-fuzzyhas no built-in counterpart: pi'sfindis glob-only, fff's is fuzzy.- pi-fff runs in its default
tools-and-uimode, so the built-ingrepandfindstay registered next toffgrepandfffind. Nothing gets replaced. fffindwith an empty pattern andpath: '**/*.ext'is the glob equivalent offind.- A headless-only
Theme not initializederror from an unrelated UI extension shows up in both conditions. It is caught and does not affect the timings.
So, does it stay in settings.json?
The tool wins are real, and on a search-heavy session at $HOME they are hard to argue with: 526 ms down to 4 ms is the difference between a search that feels instant and one that does not. The absolute numbers on smaller trees are small, a few milliseconds saved per call, so the ratio is mostly a large number dividing a small one.
What gives me pause is the recurring bill. The startup cost and the extra 17 to 51 MB are one-off; you pay them once per process. The 1,031 extra tokens ride along with every model request in the session, including the ones where no search tool is ever called. On a long tool-heavy session the searches win, and on a session that mostly talks to the model you are paying for a speedup you never use. That is the tradeoff I would want someone to show you before you add the line to your settings file, and it is the one this benchmark settled for me.