15 Questions That Expose a Weak AI Vendor Fast

Most AI demos are staged to hide the failure modes. These fifteen questions surface them before you sign, not after your residents do.

The short answer

The three fastest questions to expose a weak AI vendor are: 'Show me a live citation for that answer,' 'What happens when the agent does not know?' and 'Whose data trains the model, and does mine leave my control?' Strong vendors answer instantly with specifics. Weak ones deflect, promise a follow-up, or explain why the question does not apply.

The three questions that end most demos early

Ask these first

Ask: (1) 'Show me a citation for that answer, live, right now.' (2) 'What happens when it does not know the answer?' (3) 'Whose data trains your model and does mine ever leave my control?' A vendor who cannot answer all three cleanly in the room is not ready for your residents or your board.

Most AI demos are choreographed. The rep types the one question the system was tuned to answer, it responds fluently, and everyone nods. Your job in that room is to break the script, because your residents will break it on day one.

The uncomfortable truth: a confident wrong answer is worse than no answer. In property management, a hallucinated late fee, a made up bylaw, or an invented vendor insurance status creates liability and destroys trust with a board that already suspects the technology. The questions below are designed to force the vendor to reveal how the system behaves when it is uncertain, not just when it is showing off.

Key takeaways

  • A good AI agent cites its source on demand; a weak one asserts confidently with nothing behind it.
  • The most important behavior is what happens at the edge of knowledge, not at the center.
  • Data ownership and training rights are contractual questions, not feature questions.
  • Every serious deployment needs human approval gates and clear escalation rules.

The 15 questions, with the good answer and the red flag

Print this table and take it to the demo. The pattern to watch for is specificity. Good vendors answer with mechanisms, logs, and examples. Weak vendors answer with adjectives, roadmaps, and reassurance.

Vendor evaluation: what a strong answer sounds like versus a red flag
QuestionGood answerRed flag
Can you show a citation for that answer live?Clicks through to the exact document, page, or ticket the answer came from"We can add that" or the answer is fluent but unsourced
What happens when it does not know?It says it does not know and escalates to a named human with the context attachedIt always produces an answer; there is no 'I don't know' state
Whose data trains the model?Your data stays yours, is not used to train shared models, and is deleted on exitVague ownership language or 'we use it to improve the product'
What is your hallucination rate and how do you measure it?A number, a test method, and examples of caught errors"Our model doesn't hallucinate"
How does a human approve or override an action?Named approval gates before any outbound message or money movementThe agent acts autonomously with no gate
What does the escalation path look like?Rules by topic, time, and severity, routed to specific roles"It emails someone"
How is the agent trained on our specific communities?Ingests your docs, past tickets, and bylaws per communityOne generic model for every client
What happens to our data if we leave?Export everything, they delete their copy, confirmed in writingNo exit clause or data held hostage
Can it handle Spanish or other languages residents use?Demonstrates it live in the languages your buildings need"On the roadmap"
Show me the audit log for a real interactionFull transcript, source, timestamp, and who approved whatNo log or a summary only
How do you prevent it from inventing policy?Answers only from your knowledge base and abstains otherwise"The model is very accurate"
Who owns the agent after implementation?You own it and keep it, clear in the contractYou rent access indefinitely
What is the real cost after the first year?Transparent pricing with no surprise per-seat or per-message inflationTeaser pricing, unclear renewal
What can it not do?A specific, honest list of limits and failure modes"It can do pretty much anything"
Can I talk to a current client with my door count?Yes, a reference at similar scaleNo references or only vendor-selected quotes

The contrarian point most buyers miss: the best signal is not the demo, it is question 14. Any vendor who cannot name what their product does poorly is either lying or has never watched it fail in production. A vendor with no list of limits has never shipped at scale.

Your full demo script (checklist)

Checklist

0/12

Run this live in the room, not from the vendor's slides

If a vendor resists more than two items on this list, stop the process. In property management the cost of a wrong AI answer lands on a resident, a board, or a legal file, so the burden of proof sits with the vendor, not with you.

For the deeper technical version of the citation and abstention questions, see our guide on AI hallucinations in property management and AI agent security.

How we answer these questions about our own product

We hold ourselves to the same table. Here is how One Home Agent answers the sharpest questions, plainly, including where we draw hard lines.

Citations: Our PM ops agents answer from your own community knowledge base, not the open internet. Riley Resident and CAMeron surface the source document or past ticket behind an answer so a manager can verify it in one click.

When it does not know: The agents are built to abstain and escalate rather than guess. If Riley cannot find the answer in your documents, it says so and routes the resident to a named human with the full conversation attached. No invented policy, no invented late fees.

Whose data: The agents are trained on your communities and your data stays yours. We build the first agent free, you keep it, and your data is not used to train some shared model that helps a competitor down the street.

Human gates: Actions like outbound resident messages, violation letters, and anything touching money pass through approval rules your team sets. Victor Vendors flags an expired COI; a human decides what happens next.

The first thing we tell a property management company is what our agents will not do without a human. If a vendor is not eager to show you the guardrails, the guardrails probably do not exist.

Todd Paton, Partner, One Home Agent

Bottom line

The vendor who survives all fifteen questions is not the one with the flashiest demo. It is the one that answers 'I don't know' gracefully, cites its sources, keeps your data yours, and can name its own limits without flinching. Everything else is polish over a liability.

Put us through the fifteen questions

Bring your evaluation sheet. We will answer every question live.

We build custom AI ops agents trained on your communities, the first one free, and you keep it. Book a demo and run the whole checklist against us in real time.

See the property management agents

Frequently asked questions

Ask 'What happens when the agent does not know the answer?' A strong system abstains and escalates to a named human with context attached. A weak system always produces an answer, which means it invents policy, fees, or facts. Confident wrong answers create real liability in property management.

Sources & further reading

  1. National Association of Residential Property Managers (NARPM)
  2. Buildium Industry Research
  3. Florida DBPR, Condominiums (milestone inspections)

Keep reading

Property ManagementAI Hallucinations in Property Management: The Real Risk7 min readProperty ManagementCustom AI Agents vs Off-the-Shelf PM Chatbots8 min readProperty ManagementAI Agent Security in Property Management: The Real Test8 min read