The agency guide to vetting AI builds
Four questions to ask AI vendors before you approve a feature: guardrails, failure scenarios, rollback triggers, and when not to use AI.
—

It’s easy to buy into a vendor pitch involving AI when it includes a flashy demo and stats of saved time and money. But there’s a massive gap between a vendor promising an AI tool can work and ensuring your agency can legally and safely deploy it for a client.
When an AI build hallucinates, leaks data, or violates compliance standards, it’s your agency left explaining what happened. That’s why at Nexcess Signature Services, we collaborate directly with legal and compliance teams on every AI project we scope. They help identify the real liabilities before any contracts are signed.
If you’re evaluating an AI capability, like a chatbot, recommendation engine, or internal automation tool, here are the critical questions to guide your review process.
Does my client even need an AI build?
We’ve worked on multiple AI-based project requests, but our team has learned to ask if AI is the right answer to the problem they are trying to solve. AI hype has caused so many clients to assume AI is the answer. We suggest that all clients think through technical and regulatory considerations that affect the client, their industry, and the request. These constraints help us decide if building an AI capability is really the best solution for their needs.
If you need something predictable and consistent, a custom site block or traditional feature is often far easier to test, maintain, and explain than an AI integration.
In the end, choosing a solution that works well is far more important than its underlying technology.
How do I make sure the right constraints are in place?
To prevent AI systems from sharing unauthorized information, you need guardrails. If a vendor can’t point to specific mechanisms that stop the system from operating outside of its authority, then you should reconsider the use of this solution
Scoped retrieval is a common example. Instead of letting a chatbot generate answers from anything it was trained on, the vendor limits it to a fixed, approved set of documents, and nothing else. It narrows the risk of the AI providing misinformation, and it gives you something concrete to test and audit.
What does the failure scenario look like?
An AI failure is silent, confident, and difficult to reproduce. Unlike traditional code where the same bad input triggers the same clear error code every time, generative AI hallucinates wrong answers with total confidence. Dev teams can rarely recreate the bug because when you ask AI the same question twice, it answers with different results. Clients discover the issues only after an angry user complains.
Ask for the failure scenario in writing. What does a customer see, what does the support team see, and who gets notified?
A vendor with a real answer might describe something like this:
When the AI’s confidence score drops below a set threshold, the customer sees a fallback message (“I’m not certain about this, here’s a support article that might help” or a handoff to a live agent) instead of a guessed answer. At the same time, that interaction gets flagged and logged with the query, the response, and the confidence score attached. Support sees the flag in a queue, not buried in a support ticket, and a designated reviewer is notified if flagged interactions cross a daily threshold.
That’s a specific, testable answer. You can ask to see the confidence threshold, watch the fallback message trigger, and check the flagged-interaction log yourself.
What’s the governance plan?
Unlike traditional software, AI tools change constantly as LLMs are updated and users interact with a tool that’s built to learn. The best way to prepare for change is to have a governance plan in place before launch.
Work with the vendor and your internal team to develop a governance plan. At minimum, that means defining:
- Who’s responsible for reviewing outputs after launch
- How often reviews happen – this should be aligned with how often new information is added to the AI engine, and how quickly it is learning new information.
- Who’s authorized to use or adjust the tool
- How the outputs get used downstream
Moving from discovery to decision-making
When implementing an AI solution, it’s important to surface gaps before you get into production. We’ve turned down AI builds we were fully capable of shipping because a client’s legal team asked good questions and vendors didn’t have answers that satisfied us or our clients.
The best AI vendors to work with are the ones who lead with judgment and engage in conversation about governance, constraints, fail scenarios etc. before the build starts
If you want to learn more about our process, we created a case study that highlights how the questions above guided us away from a chatbot and toward building a rules-based system instead.
And if you’re at WordCamp 2026, Kay Lima, Jason Zinn, and I will be talking about this in more detail at our session, Your Client Wants AI. Their Lawyer Just Said No. We’ll walk through the AI build and the deterministic build side by side, plus a production checklist you won’t see in the demo.
Table of contents
Get hosting news and tips straight to your inbox
Join our community today.
Essential Hosting Resources to help your business stay ahead
Share this page