The point of this playbook
A vendor decision based on evidence, data boundaries, and exit cost-not the quality of the demo.
01
Start with the workflow
Do not evaluate an AI vendor in the abstract. Define the job, the people affected, and the current baseline before the first demo.
- What specific workflow should improve, and for whom?
- What measure would change if the product works?
- What error would be merely annoying, and what error would be unacceptable?
- Which systems and data would the product need to touch?
02
Data and model questions
The product interface is only one layer. Understand what happens underneath it and what can change without your approval.
- Is our data retained, used for training, or available to subprocessors?
- Which model providers are used, and can that provider change?
- Where is data processed and stored?
- Can administrators control retention, access, and deletion?
- What audit history is available for generated outputs and user actions?
03
Evidence and evaluation
Ask for performance on work that resembles yours. A broad accuracy claim is not evidence that the product will survive your edge cases.
- How was the product evaluated, and who chose the test set?
- Can we run a trial using representative, approved examples?
- How are hallucinations, refusals, and inconsistent outputs measured?
- What human review does the vendor assume in normal use?
- How does performance change when the underlying model changes?
04
The exit test
A tool is not truly adoptable if leaving it later would mean rebuilding the workflow from scratch.
- Can we export our data, configurations, prompts, and evaluation history?
- What happens to our data after termination?
- Which integrations or workflow steps would need to be replaced?
- What is the total switching cost after one year of use?