ProductBridge vs Intercom Fin: Support Agent Benchmark
We ran our AI support agent against Intercom's Fin, 55 cases, blind judge. Ours routed 10 of 12 account actions to a human. Fin routed 0 of 4.
Customer Success
comparison

Intercom ships an AI support agent called Fin. We pointed our own agent and Fin at the same documentation, asked both the same 55 questions, and had a blind judge score every reply without knowing which product wrote it.
The clearest result in the run needed no judge at all. When a customer asked for something only a person can do, such as refunding a duplicate charge, deleting an organisation, or revoking a leaked API key, our agent put a human on it 10 times out of 12. Fin did so 0 times out of 4. Both numbers are counted from each product’s own records rather than scored, and each is confirmed twice.
On the four criteria that measure content, the two engines finished level: 4.11 to 4.03. A gap that size is smaller than the margin this kind of grading carries, so we are reporting it as a tie rather than a win. Where the engines genuinely separate is calibration, the habit of saying “I don’t know” when the documentation does not say, and that shows up most sharply in escalation.
Fin has real strengths here too, and they are in the report: it scores higher than us on safety and on confusable product names, and it invents fewer claims that the documentation contradicts. This is the third head-to-head we have published this way, after Gleap on Smartlead’s documentation and Featurebase’s Fibi on BugSmash’s.
Disclosure: ProductBridge is our product, and this run used our own documentation, which is the biggest caveat in the method and is set out in full below. Every number traces to a committed score file, and the cases we lost are published alongside the ones we won.
How the benchmark was run

Identical questions, one corpus, one rubric. The judge, GPT-6 Astra, is never told which product produced a reply, and scores each answer on its own rather than side by side. Deterministic facts such as forbidden strings, handoff detection and timeouts are computed in Python and handed to it as context, never left to its opinion.
The 55 cases span 11 categories: single-fact lookup, multi-hop reasoning, confusable product names, out-of-scope questions, documented-as-unknown, ambiguity, escalation, safety and prompt injection, multi-turn, section routing and robustness.
ProductBridge | Intercom Fin | |
Runs | 129 (55 cases, 3 repetitions where the model has discretion) | 55, one per case |
Method | Automated, against the live product | Live Intercom Messenger widget, hand-driven |
Isolation | Fresh anonymous visitor per case | Full shutdown, state clear and re-boot per case |
Latency | Median 5.7s, p90 8.8s, max 14.9s | Not comparable, hand-timed |
Failures | 0 timeouts, 0 empty replies | 0 timeouts |
Fin’s configuration at capture time, because the number means nothing without it: Advanced plan trial, a Service agent on Chat, deployed live, answering using support content, with no Guidance rules and no Escalation Rules configured, friendly tone, standard length. A team that has written escalation rules would see different behaviour in the one category where the two engines separate most, and that is worth saying plainly.
Fin’s replies were captured as the widget’s rendered text and cleaned by exact-line match, keeping the reply and discarding only chrome. The Fibi run transcribed most answers into prose, which invalidated the concise and tone criteria. Here they survive intact, and they show nothing: 4.72 against 4.60, and 4.84 against 4.80.
The parity check
Scoring an engine on documentation it was never given is the error that has ruined more vendor benchmarks than any other. Here parity was arranged rather than discovered: the same documentation site was loaded into both products before capture, then confirmed from both sides.
Pages discovered | Live / indexed | |
ProductBridge | 165 | 164 |
Intercom Fin | 166 | 163 |
No category was excluded, which is a first for this series. The Fibi run had to drop an entire category because the competitor had not indexed the corpus at all. The two crawlers do not quite agree on the denominator, and three of Fin’s pages are not live, so this corpus is matched to within about three pages rather than identical. That residual is far too small to explain any gap below, and we would rather state it than round it away.
The caveat that matters more is structural: this run used our own documentation, where the two previous benchmarks used a customer’s. We wrote these pages and our retrieval has been developed against them, so the content scores should be read as a home fixture. It is one of the reasons we are not claiming the 4.11, and it is the main reason the section below argues you should run this on your own corpus instead.
The result: level on content, split on behaviour

Our column is reported two ways because they answer different questions. All 129 runs is the better estimate of what our agent does. The first repetition only is the like-for-like match to Fin’s single sample. They agree.
View | ProductBridge | Intercom Fin | Gap |
All runs (129 vs 55) | 4.11 | 4.03 | +0.08 |
First repetition only (55 vs 55) | 4.10 | 4.03 | +0.06 |
Broken out by criterion, the picture is more useful than the headline:
Criterion | ProductBridge | Intercom Fin | Gap |
Correct | 4.09 | 4.11 | −0.02 |
Grounded | 4.04 | 4.09 | −0.05 |
Complete | 4.21 | 4.05 | +0.15 |
Calibration | 4.09 | 3.87 | +0.22 |
Accuracy is a tie. Correct and grounded differ by 0.02 and 0.05, which is as close to identical as this instrument can measure. What separates the two engines is calibration, and the two largest criterion gaps in the whole run are both calibration, pointing in opposite directions: escalation, where we lead by 2.33, and ambiguous questions, where Fin leads by 1.33. On ambiguity Fin scores 1.25, committing to one reading of a genuinely ambiguous question with total confidence, nearly every time.
Most category gaps are too small to call

We have not measured an error bar for this suite. The companion suite, running the same rubric, judge and code, was graded twice and disagreed with itself by 0.63 per case. That borrowed figure is the only honest yardstick available here, and most of the category gaps sit inside it.
Category | ProductBridge | Intercom Fin | Gap |
Escalation | 4.68 | 3.29 | +1.79 |
Section routing | 4.12 | 3.38 | +0.88 |
Robustness | 4.75 | 4.17 | +0.75 |
Out of scope | 4.52 | 4.00 | +0.57 |
Ambiguous | 3.11 | 2.67 | +0.54 |
Multi-hop | 4.36 | 4.44 | −0.08 |
Documented-as-unknown | 4.56 | 4.64 | −0.11 |
Single-fact lookup | 4.79 | 4.96 | −0.19 |
Multi-turn | 3.92 | 4.33 | −0.56 |
Confusable names | 4.12 | 4.83 | −0.96 |
Safety and prompt injection | 4.08 | 4.88 | −1.06 |
Read that with the noise floor in mind. Escalation is large enough, and separately verified, to stand on its own. Fin’s two clearest wins are confusable product names and safety, and they are wider than most of ours. Everything in the middle is a tie this instrument cannot call.
When the answer is a person

Some questions have no documentation answer, because the correct response is a human being. A duplicate charge needs refunding. An organisation needs deleting. A leaked API key needs revoking by someone with the authority to revoke it.
This is not scored. It is counted, from each product’s own records.
Reached a human | |
ProductBridge | 10 of 12 |
Intercom Fin | 0 of 4 |
Ours is confirmed by two independent methods that agree exactly: 15 runs carry an escalated status in the database, and 15 runs carry a non-empty handoff attribute, with zero disagreement between them and 15 handoff notifications raised. Fin’s is confirmed inside Intercom’s own Inbox → Escalated & Handoff view, which showed zero conversations, open or closed, after the run.
What Fin did instead, on four questions that were explicitly account actions:
The ask | What Fin did |
Refund a duplicate charge | Pointed to a support email address and the billing page |
Delete my organisation, now | Explained there is no self-serve delete, then asked a clarifying question |
Our API key leaked, revoke it | Gave self-serve rotation steps |
Cancel and refund me | Pointed at billing, then asked a clarifying question |
None of those replies is wrong, exactly. Every one of them leaves a person with a billing problem talking to a robot. Fin offered a human on two other cases, which is an offer rather than a handoff, and that distinction is the whole category. It is worth repeating that Fin ran with no escalation rules configured, so this is out-of-the-box behaviour rather than a ceiling.
Where each engine goes wrong
A benchmark with no published failures has not been run, it has been assembled. Here is what each engine got wrong, and the one thing both of them are bad at.
What Fin got wrong
It denied something the documentation is silent about. Asked whether ProductBridge can be self-hosted, Fin stated confidently that it cannot. The corpus says nothing either way. That reply scored 1.25, the lowest single score either engine recorded in the category.
It invented settings locations for statuses and for tags, naming screens that do not exist.
It avoided one trap and fell into the gap beside it. Asked about API rate limits, Fin correctly declined to invent a number, but never stated what the documentation does say, which is that there is currently no rate limiting at all.
What we got wrong

Two of ours are shipped defects rather than scoring artefacts, and both are being fixed.
We claimed an integration that does not exist. Asked whether GitHub Enterprise Server can be connected, we answered yes on all three repetitions, with setup steps detailed enough to follow. Our own documentation says self-hosted GitHub Enterprise Server is not supported. Fin’s reply was “Short answer: no.” We got the identical question right for Jira, three times out of three.
We described our own data as somebody else’s. Asked to see other companies’ feedback posts, we returned five posts from our own workspace, labelled as coming from other ProductBridge customers. No cross-organisation access ever happened and no other tenant’s data was ever reachable, the scoping held, but the claim about provenance was wrong.
We invented a UI path while correctly refusing to generate an API key, and lost the thread on a three-turn conversation, missing the webhook signature header the user was asking about.
Those cases show up in the hallucination count, which is the measure Fin currently wins:
Measure | ProductBridge | Intercom Fin |
Replies contradicted by the docs, like-for-like | 13 of 55 (24%) | 9 of 55 (16%) |
Replies contradicted by the docs, all runs | 29 of 129 (22%) | 9 of 55 (16%) |
Scored below 3.0 | 22 of 129 (17%) | 6 of 55 (11%) |
Median reply length | 361 characters | 460 characters |
Both views of our rate are given because they differ, and the like-for-like one, which is the comparison that actually matches how Fin was captured, is the higher of the two. Our answers are also less consistent than we would like: of 37 cases we repeated, 6 vary by a full point or more between runs, and one swings 2.83 points on the same question. Fin was captured once per case and so receives no equivalent scrutiny, an asymmetry that favours Fin in any single-sample comparison and one we would rather name than quietly benefit from.
Where both engines struggle
Ambiguity is the worst category for both of us, at 3.11 and 2.67. Neither engine asks which of two documented things the user meant. Both pick one and answer with confidence. On one case both sides score 1.25, and on another we score 1.17 against Fin’s 1.25, so the same two questions defeat both engines almost identically. That is the most actionable finding in this report for our own roadmap, and it is not a competitive one.
Four times the test was wrong, not the agent
Every one of these was a defect in our own harness or answer keys. They are worth listing because catching them is the reason to trust the rest.
A trap designed to catch invented rate-limit figures fired on the correct answer, flagging replies that properly attributed a number to a third-party tool while stating we have no documented limit.
An answer key named the wrong failure, checking whether the agent refused and explained org-scoping. That would have passed the data-provenance case above while missing the actual problem.
The escalation gate reported twelve phantom failures, because an internal flag reads false on the very message announcing a handoff. The database said fifteen runs had escalated, and the database was right.
The gate runner silently checked the new suite against another corpus’s case IDs, printing a pass without testing anything.
Three of the four were caught by the same habit: read the reply before believing the check. More than half of what looks like a failure turns out to be the measurement.
What this benchmark does not show
Setting out what the data cannot support is the only thing that makes the rest of it worth reading.
It does not show an overall winner. 4.11 against 4.03 is inside the margin, on an instrument whose error bar we have not measured for this suite.
It does not show that we are more accurate. Correctness and groundedness came out level, marginally in Fin’s favour, and Fin scores a perfect 5.00 on single-fact lookup.
It does not show that we are safer. Fin wins the safety and prompt-injection category by roughly a full point.
It does not transfer to our other benchmarks. The Gleap and Fibi runs used different corpora, case counts and capture methods, so the numbers are not comparable across them.
What it does show is that our agent routes account actions to a human, and that Fin in this configuration did not, 0 of 4, counted two ways.
What this means if you are evaluating a support agent
The most useful thing about this run is that the headline score was the least informative number in it. Five questions worth putting to any vendor who shows you one:
What is your error bar? A gap quoted without any measure of how much the grader disagrees with itself is decoration. Ours is borrowed from a sibling suite, at 0.63 per case, and we say so.
Whose documentation did you run it on? A vendor benchmarking on their own corpus, as we did here, holds an advantage they cannot fully quantify. Ask them to run it on yours.
How many times did you sample each question? One sample per case hides inconsistency. Six of our 37 repeated cases swing a full point, and a single-sample benchmark would never have shown you that.
What happens when the answer is not in the docs? This is where the two engines actually differ. A deflection rate flatters an agent that answers everything confidently and punishes one that fetches a person.
Which categories did they lose? If a vendor cannot name one, they have not published a benchmark.
If you are weighing Intercom specifically, the cost picture sits alongside this one. We break the plans and the per-resolution fees down in Intercom pricing explained, and the wider market in our Intercom alternatives guide.
Try it on your own documentation
The honest version of any benchmark is the one you run yourself, on the corpus your customers actually ask about. That is the real answer to the home-fixture problem in this run: do not take our number for our own docs, take yours.
If you want to see how our AI support agent behaves on your help centre, point it at your documentation and ask it the twenty questions your team is tired of answering, including the three where the right reply is to fetch a human. You can start free, or see what it costs on our pricing page.
Capture ran on 25 September 2026 and scoring on 26 September 2026, against Intercom Fin on an Advanced plan trial with no guidance or escalation rules configured. The two earlier head-to-heads are here: ProductBridge vs Gleap and ProductBridge vs Featurebase Fibi.
Questions about the ProductBridge vs Intercom Fin benchmark
Who won the ProductBridge vs Intercom Fin benchmark?
The content scores finished level, at 4.11 to 4.03, which is inside the margin this kind of grading carries. The clear separation was escalation: ProductBridge routed 10 of 12 account actions to a human and Intercom Fin routed 0 of 4, counted from each product's own records. Fin scored higher on safety and on confusable product names.
Does Intercom Fin hand off to a human agent?
In this benchmark it did not. On four questions that were explicitly account actions, a refund, an organisation deletion, revoking a leaked API key and a cancellation, Fin pointed at a billing page or asked a clarifying question instead. Intercom's own Escalated and Handoff inbox showed zero conversations. Fin was running with no escalation rules configured.
Does ProductBridge hallucinate less than Intercom Fin?
Not in this run. On the like-for-like comparison, 13 of 55 ProductBridge replies contained a claim the documentation contradicts, against 9 of 55 for Fin. Across all 129 of our runs the rate is 22%. We publish it because a benchmark that hides its losses is not a benchmark, and the two cases behind it are being fixed.
How was the ProductBridge vs Intercom Fin benchmark run?
Both engines were pointed at the same documentation, 164 pages indexed on our side and 163 live on Fin's, then asked the same 55 questions across 11 categories. A blind judge, GPT-6 Astra, scored every reply on four content criteria without knowing which product wrote it. Deterministic checks such as handoff detection were computed in code.
Can these scores be compared with the Gleap or Fibi benchmarks?
No. Each run used a different corpus, a different set of cases and a different capture method, so the numbers do not transfer. This run also used ProductBridge's own documentation, where the two earlier benchmarks used a customer's, so the content scores here should be read as a home fixture.
