The Vendor Scorecard
The EU AI Act is about to become enforceable, procurement now has it's finished checklist, and every vertical AI vendor is about to be graded on the same seven rows.
The RFP That Arrived Pre-Written
A vertical SaaS CRO opens an enterprise RFP and hits a section that wasn’t there last year. Thirty AI questions, numbered, precise, the kind of list that only shows up when the buyer has stopped trusting the demo. Whose models run this workload, and what data leaves the tenant? Where are the audit logs, and how fast can you get them out? Who owns the output when it’s wrong? Show the opt-outs. Name the kill switch.
The thing is those questions read like the same lawyer wrote them everywhere, and in effect one did. Public AI-vendor RFP templates landed all summer. Forty-question checklists for AI platforms, thirty-question agent-procurement lists, CIO checklists formatted to copy-paste straight into a questionnaire. Procurement teams stopped inventing their diligence. They started downloading it.
And two days ago the deadline behind all of it arrived. On August 2 the EU AI Act became fully applicable, and the penalty math is not abstract…up to EUR 35 million, or 7 percent of global turnover. The buyer’s scorecard for AI vendors used to be tribal knowledge, scattered across a dozen security teams who each asked it a little differently. Not anymore. It converged, in public, into a set of rows you can see coming.
In our last article, “Trust Is the New Demo,” we argued that vertical AI deals are decided in the security review. This is the sequel procurement wrote for us…the scorecard itself, row by row, with the evidence that passes and the red flag that fails. If your next enterprise deal has not included this section yet, it is one quarter away.
What Converged, and Why This Summer
Three forcing functions stacked up this summer, and any one of them alone was easy to wave off.
Regulation Set the Floor. The AI Act’s general-application milestone passed this weekend, and the enforcement number does the talking…up to EUR 35 million or 7 percent of global turnover for the worst violations. Sell into a buyer with any European exposure and you inherit that buyer’s anxiety, no matter which country signs the contract. Compliance officers don’t read your marketing. They read the instrument, and the instrument is live now.
Standardization Gave the Anxiety a Form. ISO/IEC 42001, the AI management-system certification, is turning up in enterprise questionnaires right next to SOC 2, and it’s spreading the exact way SOC 2 did a decade ago. First a differentiator. Then a line item. Then, one renewal season, a gate. Anyone who sold software through that shift remembers how it felt. For a couple years “do you have a SOC 2” was the question that impressed a buyer when you said yes. Then it quietly became the question that ended the meeting when you said no, and nobody sent a memo announcing the switch.
The Public RFP Wave Took the Gatekeeper Out. A buyer no longer needs a mature AI-governance team to run mature AI diligence. The templates are free, the questions are numbered, and the intern doing vendor intake can grade the first pass before it ever reaches a specialist.
The money confirmed it is structural. Gartner puts AI governance platform spending at roughly $492 million for 2026, on a path past $1 billion by 2030. Grand View Research sizes the broader AI governance market at $308 million in 2025, projected toward $3.6 billion by 2033. Buyers aren’t just asking harder questions. They’re budgeting for the tooling to keep asking them, on a loop, long after the ink dries.
Separately, each of these was a headline. Stacked, they are a grading system.
The Scorecard, Row by Row
Here is what that converged scorecard actually grades, pulled from the public templates, the Act’s requirements, and the questionnaire lines now shipping in real RFPs. Seven rows. For each one…what the buyer asks, what evidence passes, and the red flag that ends the conversation.
Model Governance and Data Routing. The question: whose models run this workload, what data leaves our tenant, and what are the training opt-outs? What passes is a data-flow diagram the buyer can hand straight to their privacy team, with opt-outs documented per model provider. The red flag is a vendor who has to go ask engineering.
The Audit Trail. The question: is every AI decision logged with input, output, model version, and timestamp, and how fast can we get it? What passes is an export the buyer watches happen in minutes, live, during the demo. The red flag is “we can pull that together for you,” which is a two-week consulting engagement wearing a feature’s clothes.
Liability Assignment. The question: when the output is wrong, who owns the decision? What passes is documentation that names the owner per output type: the clinician who reviewed the recommendation, the attorney who signed off on the language, the dispatcher who confirmed the schedule. The red flag is ambiguity, because ambiguity is the thing that gets litigated.
Compliance Evidence on Demand. The question: SOC 2, ISO 42001 where you claim it, and the sector pack (HIPAA in healthcare, matter confidentiality in legal), produced as system exports. What passes downloads at 2 AM the Friday before the buyer’s audit. The red flag is a PDF a consultant wrote last spring.
Reproducibility. The question: same input, same model version, same output, yes or no, and can we test it ourselves? What passes is a documented versioning policy and a test harness the buyer can actually run. The red flag is a vendor who calls the question unfair. Regulators won’t.
Human Override and the Kill Switch. The question: how are agent identities scoped, what can a human interrupt, and what does rollback look like at 2 PM on a Tuesday when something’s on fire? What passes is a named, tested procedure with scoped permissions per agent. The red flag is an org chart where a mechanism should be.
Outcome Accountability. The question procurement learned from pricing the work itself: what did the system actually do, and what’s its error rate? What passes is a work log the buyer can reconcile against their own records. The red flag is an adoption dashboard, seats and sessions standing in for outcomes nobody measured.
If you read the trust piece, rows two through five will look familiar. That’s the trust architecture from July, now sitting inside a grading rubric. Rows one, six, and seven are what the RFP wave added, and they’re the ones where most vendors have nothing in the folder.
The Weights Are Vertical
Here’s what the public templates get wrong, and it’s exactly where the vertical operator’s edge comes back.
The rows are universal. The weights are not. A hospital system grades rows four and five the hardest, because HIPAA evidence and reproducible clinical outputs are existential and everything else is a conversation. A law firm reads row three first, liability and confidentiality routing, because privilege doesn’t survive ambiguity. A field-service or trades buyer barely looks at row five and lives on rows six and seven: can a dispatcher override it, and did the work actually get done? A property-management portfolio leans on rows two and seven at scale, because a thousand buildings throw off a thousand disputes a quarter and the logs are what settle them.
The market is already paying vendors who read the weights right. Avoca, a voice AI built for HVAC, plumbing, and field-service trades, raised more than $125 million at a $1 billion valuation in April, selling into buyers who grade almost entirely on “did the calls get handled.” EliseAI, which automates leasing and scheduling across property management and healthcare, hit a $2.2 billion valuation and now runs north of $100 million in annual revenue across most of the largest property managers in the country. Neither one cleared procurement at that scale with a generic compliance page. They cleared it because they knew which rows their vertical actually grades.
The weights change what “passing” costs, too. A horizontal vendor preps every buyer for all seven rows at full depth, which is slow and expensive. A vertical vendor ranks the rows by what their market really grades, spends hardest where the weight is, and carries lighter evidence where it isn’t. That’s not just a sales advantage. It’s a capital-efficiency advantage, and in a funding market that’s rewarding vertical AI right now, that compounds.
Same pattern as every prior piece in this series, showing up again in the diligence layer…horizontal vendors learn the rows, vertical operators know the weights.
Preparing to Be Scored
The operator response here isn’t a compliance project. It’s a GTM motion, and it fits in one quarter.
Assemble the evidence folder before the RFP shows up. One artifact per row: the data-flow diagram, the log export, the liability matrix, the compliance pack, the versioning policy, the override procedure, the work log. If a row has no artifact, that row is your roadmap, and it’s a lot cheaper to learn it from your own audit than from the debrief of a deal you already lost.
Run the scorecard on yourself every quarter, the same way you run a pipeline review. Picture a team doing it for the first time: seven rows up on a whiteboard, honest marks, and a very quiet room around row six, because nobody ever actually tested the rollback. That silence, found in August, costs a working session. Found by a buyer’s security team in November, it costs the biggest deal of the quarter.
Name an owner. The scorecard is cross-functional by nature, which in most companies is a polite way of saying nobody owns it. The vendors who pass fastest have one person who can produce any row’s evidence in a day.
Then put the passing rows into the sales motion. The scorecard is public now, so preparing for it quietly throws away half its value. Walk into the first call with the folder already open and you turn the security review from a gate into a demo, which is the whole thesis of the trust piece, now with a rubric stapled to it.
What the Scorecard Cannot Measure
The scorecard decides who is safe to buy. It does not decide who is right to buy. No row measures whether you understand the difference between a payer contract and a lease renewal, whether your roadmap tracks the vertical’s real regulatory calendar, or whether the founder picks up the phone when an implementation wobbles in week three. Relationships and domain depth still decide vertical markets. The rows just decide who gets into the final conversation.
A perfect scorecard with no standing in the vertical still loses to a trusted operator with a credible one. The failure mode to avoid isn’t a missing row. It’s treating the rows as the whole game.
The Security Review, 2027
Picture the same enterprise security review, eighteen months from now, two vendors in the calendar.
The first asks for three weeks to compile responses. The questions route to engineering, the answers route back through legal, and somewhere in week two the buyer’s champion starts apologizing to her own procurement team for the delay. That’s the moment the deal starts losing altitude.
The second one opens a folder. Seven rows, evidence per row, dated, exportable, the same folder they review internally every quarter. The security review takes forty minutes and the rest of the meeting is about the work: what the system will do, at what error rate, priced how. The buyer’s compliance officer, the person the trust piece said now decides these deals, leaves with nothing to push on.
Same product tier. Same vertical. The only difference is that one vendor treated the scorecard as an interruption and the other treated it as the sales motion. The rows were public either way.
The scorecard converged this summer. The grades start now.
The Vertical GTM Guild is where vertical SaaS operators turn shifts like this one into playbooks before the market makes them table stakes.
If this was useful:
Join the Guild community and get the weekly operator brief
Steal the Vendor Scorecard below…run it on your own company this week, one row at a time.





