Most AI procurement conversations are conducted on the vendor's terms, using the vendor's vocabulary, against a demonstration the vendor prepared. The questions below exist to change that. They are the AI vendor questions courts should ask, phrased so that a confident answer and an evasive one are distinguishable.
The test throughout is specificity. A vendor who can name the model, state where it runs, and describe what is logged is describing a system they built. A vendor who answers in terms of enterprise-grade security and industry-leading accuracy is describing a brochure.
The guide to writing a court AI policy covers what the policy should require of procurement. This article is the list.
| Question |
A good answer sounds like |
A worrying answer sounds like |
| Which model performs each function, and can you substitute it without telling us? |
Named models, version identifiers recorded against outputs, contractual notice before material change |
Reluctance to name anything, or a description of the vendor's own layer only |
| Where does processing happen, and does court content leave our environment? |
Named regions and subprocessors; cloud, on-premises, hybrid and air-gapped options |
"Enterprise-grade cloud", or their cloud as the only option |
| What was the model trained on, and is our data used to train anything? |
A written contractual commitment with a definition of what counts |
A verbal assurance, or "we don't sell your data" |
| How is every output cited back to source? |
A live demonstration on a document the vendor did not prepare |
Citation described as a roadmap item |
| What is logged when an output is generated, and can we export it? |
Tamper-resistant logs, stated retention, a working export |
Logs that exist only inside the platform |
| Is human review enforced by the product or left to our policy? |
A workflow that cannot proceed without a recorded review step |
"Our customers configure that in their own procedures" |
| What do we keep if this ends? |
Source material, outputs, logs, and any index built from court content, in a named format |
An offer to export "your documents" |
| Who else touches court data, and what is your incident history? |
Named subcontractors and third-party services, with a disclosed history |
A security posture described without naming anyone |
The reasoning behind each one follows.
The eight questions that discriminate
Which model, and can it change without notice
Ask which model or models perform each function, and whether the vendor can substitute a different one without telling you.
The reason is reproducibility. If a court relies on an output and is later asked how it was produced, an answer that includes an unspecified model of unspecified version is not much of an answer. Silent substitution also means that a system validated during procurement may not be the system running in eighteen months.
What good looks like: named models, version identifiers recorded against outputs, and a contractual obligation to notify before material changes. What should worry you: reluctance to name anything, or an answer describing only the vendor's own layer.
Where does it run, and does court data leave
Two distinct questions that are frequently answered as one.
Where processing occurs, geographically and organizationally, including any subprocessors. And whether court content leaves the court's environment at any point, including for processing that returns immediately.
For many judiciaries this is determinative rather than a preference, because the record cannot lawfully leave national infrastructure. That analysis is covered in data sovereignty and deployment choices for national judiciaries.
Ask specifically about deployment options: cloud, on-premises, hybrid, and air-gapped. A vendor offering only their cloud has answered the question even if they have not said so.
What was it trained on, and does our data train anything
Two questions again, and the second matters more.
Training data provenance affects bias and reliability, and a vendor should be able to describe it at least in general terms. But the question a court must have answered unambiguously is whether court content is used to train, tune, or improve any model, including in aggregated or de-identified form.
The answer needs to be in the contract rather than in a conversation. Court material includes sealed filings, victim details, and juvenile matters, and a general commitment not to train on customer data is worth having in writing with a definition of what counts.
How are outputs cited
Ask whether every output identifies the material it came from: document, page, timestamp, speaker where applicable.
This is not a nicety. It is what makes human review practical rather than nominal, which is what a court's own policy will require. Without citation, a reviewer must either trust the output or redo the work, and under time pressure they will do the first.
Ask to see it working on a document the vendor did not prepare.
What is logged, and can we export it
The logging questions: what is recorded when an output is generated, whether the review step is recorded, whether logs are tamper-resistant, how long they are retained, and whether the court can export them.
Exportability is the one vendors are least prepared for and courts most need, because a log that lives only inside a platform disappears with the platform. The requirement is covered further in explainability and audit trails for AI in courts.
Is human review enforced by the product or left to policy
Ask whether the workflow can proceed without a recorded review step.
A gate enforced by the product holds under time pressure. A gate that depends on staff following a policy fails exactly when it matters, which is when everyone is busy. Courts should ask for a demonstration of the workflow attempting to skip review, and observe what happens.
What do we keep if this ends
The exit questions: what the court retains, in what format, at what cost, over what period, and what happens to court data held by the vendor.
For AI systems specifically, this extends beyond the source material to the outputs, the logs, and any index or embedding built from court content. A court that can export its documents but not the audit record of how they were processed has lost the part that answers future questions. The general framing is in writing digital evidence requirements a court can procure against.
Incident history, stability, and who else is involved
Standard due diligence, applied properly: financial stability, incident and breach history, use of subcontractors and third-party services, support model and response commitments.
Ask specifically which third parties touch court data, because a vendor's own security posture is only as good as that of the services they call.
One framing point that improves the answers. Send these as written questions with space for a written response, rather than raising them in a meeting. A vendor answering in writing commits to something checkable, and the questions that produce hedged or absent written answers are precisely the ones worth pursuing. Meetings reward fluency; written answers reward having built the thing.
How VIDIZMO AI Intelligence Hub answers these
Rather than restating capabilities, the useful thing is where the answers sit.
Model control: customer-controlled models, so a court determines which model runs rather than inheriting the vendor's choice. Hosting: cloud, on-premises, hybrid, and air-gapped deployment, so court content need not leave the institution. Citation: source citations attached to retrieved and generated content, identifying document, page, and timestamp. Logging: audit records covering processing and review, retained and exportable. Review enforcement: human-review and approval gates configured as workflow steps rather than policy statements.
Where a court should still press: exercise the export during evaluation rather than accepting a description, and put the training commitment in the contract rather than in an email.
Using the list
Send the questions in writing before the demonstration rather than after, and score the written answers.
Demonstrations are designed to be impressive and are a poor discriminator. Written answers to specific questions are a good one, because the vendors who can answer them precisely are a different set from the vendors who demonstrate well.
Request a demo once you have the written answers, so the demonstration tests what the answers claimed.