15 Questions That Expose a Weak AI Vendor Fast
Most AI demos are staged to hide the failure modes. These fifteen questions surface them before you sign, not after your residents do.
The short answer
The three fastest questions to expose a weak AI vendor are: 'Show me a live citation for that answer,' 'What happens when the agent does not know?' and 'Whose data trains the model, and does mine leave my control?' Strong vendors answer instantly with specifics. Weak ones deflect, promise a follow-up, or explain why the question does not apply.
The three questions that end most demos early
Ask these first
Ask: (1) 'Show me a citation for that answer, live, right now.' (2) 'What happens when it does not know the answer?' (3) 'Whose data trains your model and does mine ever leave my control?' A vendor who cannot answer all three cleanly in the room is not ready for your residents or your board.
Most AI demos are choreographed. The rep types the one question the system was tuned to answer, it responds fluently, and everyone nods. Your job in that room is to break the script, because your residents will break it on day one.
The uncomfortable truth: a confident wrong answer is worse than no answer. In property management, a hallucinated late fee, a made up bylaw, or an invented vendor insurance status creates liability and destroys trust with a board that already suspects the technology. The questions below are designed to force the vendor to reveal how the system behaves when it is uncertain, not just when it is showing off.
Key takeaways
- A good AI agent cites its source on demand; a weak one asserts confidently with nothing behind it.
- The most important behavior is what happens at the edge of knowledge, not at the center.
- Data ownership and training rights are contractual questions, not feature questions.
- Every serious deployment needs human approval gates and clear escalation rules.
The 15 questions, with the good answer and the red flag
Print this table and take it to the demo. The pattern to watch for is specificity. Good vendors answer with mechanisms, logs, and examples. Weak vendors answer with adjectives, roadmaps, and reassurance.
| Question | Good answer | Red flag |
|---|---|---|
| Can you show a citation for that answer live? | Clicks through to the exact document, page, or ticket the answer came from | "We can add that" or the answer is fluent but unsourced |
| What happens when it does not know? | It says it does not know and escalates to a named human with the context attached | It always produces an answer; there is no 'I don't know' state |
| Whose data trains the model? | Your data stays yours, is not used to train shared models, and is deleted on exit | Vague ownership language or 'we use it to improve the product' |
| What is your hallucination rate and how do you measure it? | A number, a test method, and examples of caught errors | "Our model doesn't hallucinate" |
| How does a human approve or override an action? | Named approval gates before any outbound message or money movement | The agent acts autonomously with no gate |
| What does the escalation path look like? | Rules by topic, time, and severity, routed to specific roles | "It emails someone" |
| How is the agent trained on our specific communities? | Ingests your docs, past tickets, and bylaws per community | One generic model for every client |
| What happens to our data if we leave? | Export everything, they delete their copy, confirmed in writing | No exit clause or data held hostage |
| Can it handle Spanish or other languages residents use? | Demonstrates it live in the languages your buildings need | "On the roadmap" |
| Show me the audit log for a real interaction | Full transcript, source, timestamp, and who approved what | No log or a summary only |
| How do you prevent it from inventing policy? | Answers only from your knowledge base and abstains otherwise | "The model is very accurate" |
| Who owns the agent after implementation? | You own it and keep it, clear in the contract | You rent access indefinitely |
| What is the real cost after the first year? | Transparent pricing with no surprise per-seat or per-message inflation | Teaser pricing, unclear renewal |
| What can it not do? | A specific, honest list of limits and failure modes | "It can do pretty much anything" |
| Can I talk to a current client with my door count? | Yes, a reference at similar scale | No references or only vendor-selected quotes |
The contrarian point most buyers miss: the best signal is not the demo, it is question 14. Any vendor who cannot name what their product does poorly is either lying or has never watched it fail in production. A vendor with no list of limits has never shipped at scale.
Your full demo script (checklist)
Checklist
0/12Run this live in the room, not from the vendor's slides
If a vendor resists more than two items on this list, stop the process. In property management the cost of a wrong AI answer lands on a resident, a board, or a legal file, so the burden of proof sits with the vendor, not with you.
For the deeper technical version of the citation and abstention questions, see our guide on AI hallucinations in property management and AI agent security.
How we answer these questions about our own product
We hold ourselves to the same table. Here is how One Home Agent answers the sharpest questions, plainly, including where we draw hard lines.
Citations: Our PM ops agents answer from your own community knowledge base, not the open internet. Riley Resident and CAMeron surface the source document or past ticket behind an answer so a manager can verify it in one click.
When it does not know: The agents are built to abstain and escalate rather than guess. If Riley cannot find the answer in your documents, it says so and routes the resident to a named human with the full conversation attached. No invented policy, no invented late fees.
Whose data: The agents are trained on your communities and your data stays yours. We build the first agent free, you keep it, and your data is not used to train some shared model that helps a competitor down the street.
Human gates: Actions like outbound resident messages, violation letters, and anything touching money pass through approval rules your team sets. Victor Vendors flags an expired COI; a human decides what happens next.
“The first thing we tell a property management company is what our agents will not do without a human. If a vendor is not eager to show you the guardrails, the guardrails probably do not exist.”
Todd Paton, Partner, One Home Agent
Bottom line
The vendor who survives all fifteen questions is not the one with the flashiest demo. It is the one that answers 'I don't know' gracefully, cites its sources, keeps your data yours, and can name its own limits without flinching. Everything else is polish over a liability.
Put us through the fifteen questions
Bring your evaluation sheet. We will answer every question live.
We build custom AI ops agents trained on your communities, the first one free, and you keep it. Book a demo and run the whole checklist against us in real time.
See the property management agentsFrequently asked questions
Ask 'What happens when the agent does not know the answer?' A strong system abstains and escalates to a named human with context attached. A weak system always produces an answer, which means it invents policy, fees, or facts. Confident wrong answers create real liability in property management.
Sources & further reading