Most corporate AI evaluations open with the same two questions – how good is the model, and what features does it have.
Both are reasonable questions. However, both will matter less than expected within eighteen months, because model quality is converging and feature lists converge with it. That convergence is already showing up at the infrastructure level, not just the model level. The questions that predict whether a system still works in production a year from now are about governance. And these questions are rarely asked in a first vendor call.
These are six questions worth asking. They apply to any vendor. The point is not who you choose, it is the principle behind the choice.
Which sources will it treat as authoritative, and who decides
Every vendor answers yes to "can it connect to our systems." That question sorts nobody.
The useful version is what happens when the same figure exists in three places and the numbers do not match.
Which one wins? Who sets that hierarchy, and can it differ by question type, since the authoritative source for a headcount number is rarely the authoritative source for a contract term.
Then the harder follow-up: what happens when the authoritative source is out of date?
A system that always defers to the designated source of truth will confidently repeat a stale figure.
A system that always prefers the most recent file will confidently prefer an unapproved draft. Ask which failure mode the vendor has chosen, because they have chosen one.
How are permissions enforced, and at what level
Ask whether permissions are inherited from existing systems or defined independently inside the new tool. Inherited is faster to deploy, but it inherits every gap in the existing setup. Independent is more precise and more work to maintain.
Then ask the question most vendors are not ready for – do permissions apply to documents or to answers?
These are not the same thing. Someone may be cleared to open a dataset and still not be the right recipient of a synthesized answer drawn from it, because the synthesis reveals a pattern the raw rows did not. Compensation data is the obvious case. Pipeline data and incident history are less obvious and just as sensitive.
A vendor who has thought about this will have a clear answer. A vendor who has not will hear the question as a rephrasing of the first one.
What does it do when it should not answer
Every system answers. But far fewer ones are built to decline.
Ask what triggers a hold. Ask whether the system can return a structured answer that surfaces a conflict rather than resolving it, since sometimes the conflict is the thing that the business needs to see. Ask what the escalation path to a human looks like, who receives it, and whether the requester is told why.
The instinct is to treat refusal as a limitation. In reality, it is the clearest signal that a vendor has deployed into environments where a wrong answer had consequences.
Can it show where an answer came from
Ask whether lineage is attached to every answer by default or produced on request. Default is meaningfully better, because nobody requests lineage until the answer is already being challenged, and by then the context is gone.
Ask whether it shows what was excluded, not only what was used. An answer that says which source it drew on is useful. An answer that says which source it deliberately set aside, and why, is the one that holds up in a room.
Gartner projects that 60 percent of AI initiatives will fail due to inadequate data management practices, which is precisely the gap a source line is designed to close.
Then ask how long that record persists. Lineage that expires after thirty days is not lineage; it is only a log.
What happens when the underlying model changes
Models get replaced. The one in the contract will not be the one running in two years.
So ask what happens to everything the company built on top. The source hierarchy, the permission rules, the workflow connections, the accumulated decisions about what counts as authoritative. Is that portable, or is it tied to one model, one vendor, one architecture?
This is the lock-in question and it is the one most often skipped, because it describes a problem that does not exist yet. The context a company builds is the part it owns, and it is worth confirming that it stays owned.
Who owns this internally, and what does the vendor expect from us
This question reveals implementation reality faster than any demo.
Ask how much of the work is the company's rather than the vendor's. Ask how long before the first useful answer, not the first working deployment. Ask what happens if the internal owner leaves halfway through.
Vendors who have done this before answer specifically, usually with a number and a caveat.
Vendors who have not answer in principles. The difference is audible.
What good answers sound like
The six questions matter less than the pattern in how they are answered.
Specific beats confident. A vendor who says "that part is genuinely hard, here is how we handle it and here is where it still breaks" is more credible than one who reports everything solved. The first has deployed something. The second has a slide.
Corporate AI is not a product category where the strongest demo wins. It is one where the system that knows its own limits is the one still trusted a year later. The questions above are worth asking because the answers reveal which kind you are buying.

