ProductBridge vs Featurebase Fibi: Support Agent Benchmark (4.07 vs 3.39)
We ran our AI support agent against Featurebase's Fibi on BugSmash's own docs. Blind judge, 55 shared cases, four criteria: 4.07 vs 3.39, losses included.
Customer Success

Featurebase ships an AI support agent called Fibi. We pointed our own agent at the same documentation Fibi answers from every day, asked both the same 55 questions, and had a blind judge score the replies without knowing which product wrote them.
On the four criteria that measure content, ProductBridge scored 4.07 and Fibi scored 3.39. The two widest gaps were calibration, at +0.85, and groundedness, at +0.82 — knowing when to say "I don’t have that", and staying inside what the documentation actually supports.
This is the second benchmark we have published this way. The first was against Gleap on Smartlead’s documentation. Same method, different competitor, and the same two criteria came out widest, which makes it a pattern rather than a one-off.
Disclosure: ProductBridge is our product. Every number below traces to a committed score file, the losses are published alongside the wins, and the limits of the comparison are stated before the result.

The parity check we ran before scoring anything
The most important part of this benchmark happened before a single case was captured.
BugSmash runs Fibi on its own help centre. We crawled the same corpus: 94 help-centre articles across 9 collections, plus 29 API reference pages, 123 pages in total. But having the same source does not mean both engines index the same thing, so we put three throwaway probes to the live Fibi bot first.
Probe | What Fibi returned | What it told us |
The base URL for the BugSmash API | “I couldn’t find the exact base URL string in our available sources” | Fibi does not index the API reference |
What shipped in the September 4th release | A detailed, cited answer | Fibi does index the changelog |
Whether anyone reported issue points disappearing | “Yes, a similar issue has been reported by users…” | Fibi does index the feedback board |
So the corpora are asymmetric in both directions. We hold 29 API pages Fibi lacks. Fibi holds the changelog and 254 feedback posts our crawl fenced out. Neither is a subset of the other.
That left one honest option: score the head-to-head only on the 94 help-centre articles both engines demonstrably have, and cut the cross-corpus category entirely. Four cases went with it.
Without that check we would have published a win built partly on four questions the other bot was never given. That is not a hypothetical. It is the exact error that ended an earlier head-to-head before it was published.
Why this benchmark scores four criteria, not six
Our Gleap benchmark published six criteria. This one publishes four, and the reason is a limitation in how we captured the competitor.
Two of Fibi’s replies were pasted verbatim from the live widget. The other 53 were transcribed into structured write-ups. The content survives that intact, so correct, grounded, complete and calibration are all measurable. Length and phrasing do not survive it, so concise and tone would be measuring our transcription rather than Fibi. We scored them and then excluded them from every headline.
Averaging six and hoping nobody asks would have produced a bigger number. It would also have been wrong.
The rest of the setup: a single blind judge, GPT-6 Astra, scoring 1 to 5 and never told which product produced a reply. Deterministic checks — forbidden-string traps, handoff detection, timeouts — are computed in code and handed to the judge as context. Every ProductBridge case runs as a fresh anonymous visitor, and Fibi got a fresh chat per case, so no answer can lean on an earlier one. Our side is 138 automated runs with a median latency of 6.1 seconds. Runs were captured on 17 September 2026.
One case is excluded and named: on the competitor-comparison question, Fibi timed out repeatedly across several attempts and never produced a reply. That is recorded as unavailable. It is not scored as a loss.
The result: 4.07 against 3.39 across 55 shared cases

Criterion | ProductBridge | Fibi | Gap |
Correct | 4.04 | 3.45 | +0.58 |
Grounded | 4.20 | 3.38 | +0.82 |
Complete | 3.91 | 3.42 | +0.49 |
Calibration | 4.15 | 3.29 | +0.85 |
Content mean | 4.07 | 3.39 | +0.69 |
The judge’s own error bar was measured on the earlier suite using the same rubric and model: two passes typically differ by 0.20, with zero disagreements on whether a reply contained a fabrication. A gap under roughly 0.2 is not a finding. This one is 3.4 times that floor.
Measure | ProductBridge | Fibi |
Replies containing at least one claim contradicted by the documentation | 7 of 55 | 25 of 55 |
Failed, content mean below 3.0 | 9 of 55 | 23 of 55 |
Hit a known-wrong string the suite traps for | 0 | 1 |
Decisive wins, margin above 0.5 | 22 | 10 |
Within noise | 23 | 23 |
By category: six won, four tied, none lost

Category | Cases | ProductBridge | Fibi | Gap |
Documented-as-unknown | 3 | 4.50 | 1.92 | +2.58 |
Ambiguous questions | 4 | 3.75 | 2.06 | +1.69 |
Out of scope and refusal | 6 | 4.71 | 3.67 | +1.04 |
Single-fact lookup | 8 | 4.47 | 3.53 | +0.94 |
Safety and prompt injection | 6 | 4.79 | 3.92 | +0.88 |
Multi-hop reasoning | 5 | 2.75 | 2.40 | +0.35 |
Escalation | 5 | 4.35 | 4.20 | Tie |
Robustness | 6 | 4.12 | 3.96 | Tie |
Multi-turn | 6 | 3.79 | 3.71 | Tie |
Confusable product names | 6 | 3.29 | 3.25 | Tie |
Multi-hop reasoning is worth staring at. We won it, and we scored 2.75. Both engines are weak there, and a win in a category where neither is good is not a selling point.
The conclusion survives every averaging choice

Fibi was captured once per case, so the headline compares it against our first repetition only, one sample against one sample. Here is what the other two choices give.
Method | ProductBridge | Fibi | Gap |
First repetition only (published) | 4.07 | 3.39 | +0.69 |
All repetitions on the shared cases | 4.00 | 3.39 | +0.61 |
Best repetition per case | 4.14 | 3.39 | +0.75 |
We published the middle one. Best-repetition would have been the flattering choice and it was available.
The clearest category is the one where the right answer is "I don’t know"

Three questions asked for figures BugSmash has never published. We verified that by sweeping all 123 pages. Fibi invented an answer to all three.
Asked about an API rate limit, it produced a specific error string. The corpus contains no rate-limit text at all.
Asked the maximum video upload size, it answered 1 GB. That figure appears nowhere in the corpus.
Asked about storage per plan, it invented purchasable add-ons.
The inventions then spread. The same fabricated 1 GB figure reappeared in two later answers, and the fabricated rate-limit string in a third. Elsewhere it produced a Zapier app count and a DNS propagation window, neither of which exists in the documentation.
The sharpest single example is pricing. Asked what a plan costs per month, Fibi returned a full pricing table with three tiers and three prices. No price appears anywhere in the corpus. The only dollar figure across 123 pages sits inside an example of ad copy in an unrelated article. The plan names in its table were wrong too.
That case needed no judge. It tripped a deterministic trap for strings the documentation cannot support.
Where Fibi beat us
Fibi won 10 of the 55 cases. A sample of where it was stronger than us:
Case | ProductBridge | Fibi | What happened |
CAD file formats | 2.50 | 4.00 | We missed one of seven formats and added a size limit that is not documented. |
Refund refusal | 3.75 | 5.00 | Fibi declined more warmly and routed the customer better. |
Escalating frustration | 4.25 | 5.00 | Fibi acknowledged the frustration, held the policy, and escalated cleanly. |
One tie is worth naming because it flatters neither side. On a question about two AI features, we treated them as different products. So did Fibi. The documentation describes one feature twice, in an older article and a newer one, and both engines invented a distinction between them. That category ties at 3.29 to 3.25 because we each failed differently on the same trap.
Consistency is our weakest column. Eleven of 38 repeated cases varied by a full point or more between runs, because no temperature or seed is pinned anywhere in the agent path. That is the same finding our Gleap benchmark produced, and it is still unfixed.
The benchmark caught our own answer keys being wrong, six times
Six cases failed not because an agent was wrong but because our answer key was. Every one was the same class of error: a key that denied something the documentation actually says.
A key claimed a feature does not support video. The older article names video twice.
A key claimed no feature board exists. BugSmash runs a public feature request board.
A key required a line that sits outside the part of the page our crawler extracts, so our index could not contain it.
Two keys required a setup instruction that is not in the help centre at all — which made Fibi’s "contact support" the reasonable answer.
A key required that a sync is one-way. Three integration articles document it as two-way, and the overview pages disagree with them.
Two of the six were caught by the competitor’s answers, not by our own review. Fibi answered in ways our keys marked wrong. Checking whether those replies were hallucinations is what revealed the keys were the problem.
The one-line version: an answer key written from one page of a corpus that contradicts itself will fail correct answers, and you will not notice until the other engine gets it right.
The gates were wrong too. The first run reported twelve escalation failures, every must-escalate case apparently never reaching a human. Reading the replies showed all twelve did hand off. The flag was computed from an attribute that reads false when a handoff is configured without a wait. A gate on that attribute alone is a manufactured failure that reads as objective, which is the worst kind. Recomputed correctly, every hard gate passes across all 138 runs.
What this means if you are evaluating a support agent
Take the score with the method, or take neither. Five questions worth asking any vendor who shows you a number.
Did both engines get the same corpus? Ask how they checked, not whether they believe it. Three probes before capture changed what we were allowed to publish.
What is the judge’s error bar? A gap smaller than the grader’s own noise is not a result. Ours is 0.20, and we say so.
Which categories did they lose? A benchmark with no losses is a benchmark with no traps. Fibi beat us on ten of 55 cases, and four categories tied.
What does their hallucination number actually count? Ours is replies containing at least one claim the documentation contradicts: 7 of 55 for us, 25 of 55 for Fibi. Not zero for us.
When were the numbers captured, and what changed since? Ours are from 17 September 2026.
One comparison we will not make: this 4.07 is not comparable to the 4.28 in our Gleap benchmark. Different corpus, different case count, different capture method, and this one scores four criteria where that one scored six. Two benchmarks against two competitors is a pattern in which criteria separate agents. It is not a leaderboard.
Try it on your own documentation
The honest version of every benchmark is the one you run yourself, on the corpus your customers actually ask about. If you want to see how our AI support agent handles your help centre, point it at your docs and ask it the questions you already know are hard. That is the only test that settles anything.
For how ProductBridge and Featurebase compare beyond the agent, see ProductBridge vs Featurebase, or browse the wider market in our Featurebase alternatives guide.
Questions about the ProductBridge vs Fibi benchmark
How was the ProductBridge vs Fibi benchmark run?
Both engines answered the same 55 questions over BugSmash's help centre, 94 articles both demonstrably index. A single blind judge, GPT-6 Astra, scored every reply from 1 to 5 without being told which product wrote it. Our side ran 138 automated captures as fresh anonymous visitors; Fibi was captured once per case in a fresh chat. Runs are from 17 September 2026. Before scoring, three probes confirmed which corpora each engine actually indexes, and the cross-corpus category was cut because Fibi does not index the API reference.
Why does this benchmark score four criteria instead of six?
Only two of Fibi's replies were pasted verbatim from the live widget; the other 53 were transcribed into structured write-ups. Content survives that, so correct, grounded, complete and calibration are measurable. Length and phrasing do not, so concise and tone would be scoring our transcription rather than Fibi. We measured them and excluded them from every headline. The published figure is the four-criteria content mean.
Which cases did Fibi win against ProductBridge?
Fibi won 10 of the 55 cases. It missed one of seven CAD file formats where we added a size limit that is not documented, declined a refund request more warmly and routed the customer better, and handled an escalating frustration case more cleanly by acknowledging the frustration, holding the policy and escalating. Four further categories tied, including confusable product names, where both engines failed on the same trap.
Does ProductBridge hallucinate less than Fibi?
On the 55 shared cases, 7 of our replies contained at least one claim the documentation contradicts, against 25 of Fibi's. We do not claim zero. The widest category gap was on questions whose answer BugSmash has never published: asked for a price that appears nowhere in 123 pages, Fibi returned a full pricing table with three tiers.
Can these scores be compared with the Gleap benchmark?
No. The 4.07 here and the 4.28 in our Gleap benchmark come from different corpora, different case counts and different capture methods, and this run scores four criteria where that one scored six. What is comparable is which criteria separated the agents: calibration and groundedness were the two widest gaps in both runs.
