Buying question, answered

Which AI SDR vendors publish evidence for their reply rate and meeting claims?

AI SDR agents are sold on replies and meetings booked. The GTM Tech Index grades all 71 vendors in the lane on whether that claim has a measurement behind it, and the answer is the sharpest gap in the category: 51 of 71 can prove the AI is real, and exactly one publishes its outcome numbers with the full measurement basis stated.

The answer: who publishes evidence a buyer can check

With the full basis stated, sample, timeframe and definitions: Tuco AI A (AI Centrality C). With real published evidence short of the full basis: Alta B, Amplemarket B, Clay B, Coldreach B, Cooby B, Evergrowth B, Kronologic B, LeadNitro B, Leadspicker B, MarketBetter B, Octave B, Playbook AI B, Rasayel B, Reply.io B, SuperAGI B. Every name links to the grade record and its sources.

The capability everyone can prove, the outcome almost nobody measures

51 of 71

document AI Centrality at a checkable standard: the AI is real, and they can show it. The best documented axis in the lane.

16 of 71

document Operational and Outcome Evidence: the number the product is bought on, with 1 carrying the full measurement basis.

The census underneath: 1 at A, 15 at B, 34 at C, meaning headline percentages with no stated basis or logos standing in for results, and 21 at D, meaning no outcome evidence beyond assertion on a product sold on its results. For scale, the whole index documents this axis at 30 percent, so the lane most aggressively sold on outcomes sits at 23 percent, below the market it is promising to outperform.

The questions buyers actually ask

Which AI SDR vendors publish evidence for their reply rate and meeting claims?

Of the 71 vendors in the AI SDR lane of the GTM Tech Index, exactly one publishes measured outcomes with the full basis stated, sample, timeframe and metric definitions: Tuco AI. Another 15 publish real outcome evidence, with named customers and numbers, where the measurement basis is incomplete: Alta, Amplemarket, Clay, Coldreach, Cooby, Evergrowth, Kronologic, LeadNitro, Leadspicker, MarketBetter, Octave, Playbook AI, Rasayel, Reply.io, SuperAGI. The rest of the lane is the finding. 34 vendors publish headline percentages with no stated basis or let customer logos stand in for results, and 21 publish no outcome evidence beyond assertion, on products sold specifically on their outcomes. Every named vendor links to its grade record and the sources behind it.

What separates evidence from a marketing number?

The basis. A reply rate is a fraction, and a fraction with no denominator is a decoration. The three questions that turn a claim into a measurement: across what population, meaning whose lists, what industries, what seniority of target; over what period, because a launch month is not a year; and under what definition, because a reply that says stop counting me is technically a reply. The strongest thing a vendor can publish is the number with all three stated. The question to send any AI SDR vendor before buying: what is your reply rate or meeting rate, across what population and what period, rather than emails sent? A vendor that has measured will answer in a paragraph. A vendor that has not will answer with a case study.

Why do so few AI SDR vendors publish measured results?

Two reasons, one fair and one not. The fair one: outbound results genuinely vary with list quality, offer and market, so a universal reply rate is close to unpublishable in good conscience, and a careful vendor may decline to commit to one. The index does not penalise that caution by itself. What the axis grades is different: whether the claims a vendor does choose to publish carry their basis. A vendor with no numbers and no claims sits differently from one advertising 3x more meetings with nothing behind it. The less fair reason is that an unmeasured claim is unfalsifiable, and in a lane where 51 of 71 vendors can prove the AI is real, only 16 can show the outcome is, which suggests the measurement gap is a choice the market has made together.

How does the GTM Tech Index grade Operational and Outcome Evidence?

On whether a published claim can be checked. Measured outcomes published with their basis: sample, timeframe, and metric definitions stated, so a buyer can tell a measurement from a marketing number. A C means: Outcome claims are headline percentages with no stated basis, or customer logos standing in for results. A D means: No outcome evidence published beyond assertion, on a product sold on its results. Logos are not evidence and prestige is not measurement: a logo proves a contract was signed, not that anything worked, and funding rounds and quadrant placements do not move this grade at all. Every grade carries a source basis and a verification date on the vendor profile, and vendor reported statistics are recorded as vendor reported, never restated as independent results.

Source basis and limits

Lane membership is primary or secondary, the same rule the per lane pages use. Grades record what a vendor has published, not what its product achieves: a careful vendor can decline to promise a universal reply rate and still grade well by publishing the basis behind the claims it does make. A grade belongs to the vendor, and on cross listed vendors it may have been earned on a different product line, so read the notes on the profile before relying on a single letter. Figures regenerate from the live index with npm run sdr-evidence-stats and were last built 2026-09-02. The lane's full shortlist across its deciding axes is at best AI SDR agents, what models sit under these products at model transparency, who publishes pricing at pricing transparency, and the standards on the methodology page.

Contact us

Found a vendor we missed? Have feedback on the index? We’d love to hear from you.