AI for Insurance Agents: Ten OpenAI Capabilities Across a Twenty-Brand CRM
Insurance agents selling Medicare, senior health, life, and final expense work inside a CRM that serves about twenty white-labeled brands from one codebase. Teamvoy built the AI layer: ten capabilities on OpenAI, from an unattended document parser to a thirteen-tool agent, each enabled tenant by tenant. Wondering how AI reaches agents in a regulated product where twenty brands share a deployment?
01. Summary
The platform runs a multi-tenant CRM for insurance agents and agencies. One Rails 8.1 backend and one Next.js 16 agent-facing application serve about twenty white-labeled brands, each with its own domain, credentials and tenant isolation. OpenAI is the only model provider in the stack, with no second vendor and no self-hosted model. Retrieval runs on pgvector inside the tenant Postgres rather than an external vector database, so embedded customer data sits inside the same isolation boundary as the records it describes – which is the condition for touching health and Medicare information at all.

— 10 weeks to the first production tenant, 10 AI capabilities shipped, 20 white-labeled brands on one codebase.
Work started 2025-08-21 and the first production tenant went live in October 2025. Ten capabilities have shipped since: eight in production, one in limited production on a single tenant, one in pilot on internal environments.
| PROJECT DETAIL | DESCRIPTION |
|---|---|
| Client | A US insurance distribution platform |
| Industry | Insurance – Medicare, senior health, life, final expense |
| Status | Live in production |
| Tenancy | One codebase, ~20 white-labeled brands, isolated per tenant |
| Technology | Rails 8.1, Next.js 16, OpenAI, pgvector on Postgres, Python FastAPI, LangGraph |
| Team and timeline | First commit 2025-08-21, first production tenant October 2025 |
Example: where AI meets an agent’s day
| AGENT MOMENT | WHAT THE AGENT DID BEFORE | WHAT HAPPENS NOW |
|---|---|---|
| A job application lands | Open the email, read the resume, hand-key a contact | The record is already in the CRM |
| Find a segment of clients | Learn the advanced filter model, or avoid it | Type the segment as a sentence |
| Prepare for a call | Scan activity, notes and email history | Read a briefing card, then ask follow-ups |
| Write a marketing email | Draft from scratch, wait on review | Draft in seconds, administrator approves before send |
| Ask how a feature works | Open a ticket | Ask the assistant, answered from this brand’s own docs |
02. Problem
AI was obvious for every one of these moments. Twenty brands on one deployment made it hard.
Agents were hand-keying applicants, avoiding the product’s deepest search surface, and reading long client histories before every call. Each is a clear candidate for a model. The difficulty lay beneath the features: tenants store health and Medicare information, unreviewed marketing copy reaching a consumer is a regulatory problem rather than a quality problem, and twenty brands share a codebase but not a feature set, so a generic answer is often the wrong answer.
- Keep embedded customer data inside the tenant boundary, which ruled out an external vector store.
- Put a person between AI copy and a consumer, structurally rather than by policy.
- Let tenants differ in features, terminology and risk appetite without a release per tenant.
- Make AI behavior queryable after the fact, per tool and per tenant.
03. Solution
Ten capabilities, one provider, and a configuration layer that carries the differences between brands
Each capability shipped as a service behind per-tenant predicates, then the parts that differ between brands – system prompt, available tools, model – moved into a scope-configuration table keyed by tenant, role, and individual user. Model selection is per workload rather than uniform: reasoning capacity is reserved for the agent loop, and high-volume paths such as classification, briefing generation and retrieval synthesis run on gpt-4o-mini under strict schemas.

— Ten capabilities, three maturity levels, one model provider. Wherever output feeds code rather than a person, it is parsed as structure rather than prose.
What does an agent actually get?
Five of the ten capabilities sit directly in the daily path. Contact records open with a short briefing card generated by relationship stage, so a cold lead and a renewing client produce different summaries. Natural-language search turns a typed sentence into a structured filter that the existing advanced-search engine runs unchanged. A conversational briefing takes follow-up questions across a client’s activity, notes, and email history. Product help answers from the tenant’s own documentation. Marketing email is drafted, then queued for approval.
How does retrieval stay inside the tenant boundary?
Queries are embedded and matched against chunked corpora held as pgvector embeddings in the tenant Postgres, using cosine distance. Documentation chunks are scoped per tenant, so a white-labeled brand retrieves only its own material and cannot surface another brand’s features or terminology. The same retrieval pattern selects fields for natural-language search, which keeps a large search schema out of the prompt.
Why replace intent routing with an agent?
The first assistant asked users to pick an intent from a menu, which pushed the internal architecture onto the agent. A classifier removed the menu. The classifier then hit its ceiling: every new capability meant a new route and a new branch, and a request spanning two capabilities couldn’t be served. A ReAct-style orchestrator on gpt-5.4 replaced it, planning across thirteen tools and taking actions rather than answering, with the loop bounded at ten iterations and a twenty-message history window.

— Thirteen tools behind one text box, including an explicit clarification tool for when the agent should ask rather than guess.
How does one codebase behave like twenty products?
Through the scope-configuration table. The system prompt, the tool set and the model can each be overridden per tenant, per role or per individual user, so withdrawing a tool from one brand or correcting a prompt is a row update rather than a deploy. This is also what makes staged rollout practical: every AI feature walks development, staging, demo, a single production tenant, then broader enablement, narrowed during initial validation to a short list of named administrator accounts.
What stops one brand’s data reaching another?
Input is length-capped and screened for prompt-injection patterns. Tool calls are validated against a per-scope allowlist before execution, so a model that hallucinates a tool it should not have cannot reach it. Tool results are truncated before entering context. Model output is scanned for the names of the other nineteen tenant brands and redacted on match. In a multi-tenant deployment, the worst realistic failure is not a wrong answer; it is a right answer about someone else’s business.
What is still being hardened?
The evaluation replay runs on demand rather than as a blocking CI gate against a versioned golden dataset, and promoting it is the top item on the list. Some tenant rollout predicates still live in Ruby and need a deploy to change, where they belong in the configuration store. Observability is error-centric, with no distributed tracing around the agent loop. Per-token cost metering is wired but not enabled, and the RSpec suite is advisory in CI rather than blocking. These are sequenced rather than unknown.
| TECHNOLOGY | ROLE IN THE SOLUTION |
|---|---|
| gpt-5.4 | Plans and acts across thirteen tools in the agent loop |
| gpt-4.1 | Email generation and refinement, agent bio rewrite |
| gpt-4o / gpt-4o-2024-08-06 | Natural-language search, conversational briefing, document parsing |
| gpt-4o-mini | Intent classification, briefing cards, retrieval synthesis |
| pgvector on tenant Postgres | Retrieval inside the isolation boundary, with no external vector store |
| Rails 8.1, Next.js 16 | Backend services and the agent-facing application |
| Python FastAPI, LangGraph | Pilot runtime for workflow generation, proxied for entitlement gating |
| Sentry, PostHog | Per-tenant error environments, product analytics |
"Twenty brands differ in features, terminology and risk appetite. Without per-tenant control of the prompt and the tool set, we would have been shipping code per tenant, and the rollout would have stopped at three."
04. Results
Ten capabilities live across about twenty brands, with tenant differences carried in configuration rather than code
Eight capabilities run in production, one in limited production, one in pilot. Every AI action is a durable record with the model used, token counts, tool arguments, status, duration in milliseconds and error text, plus per-message feedback. AI spend is metered against user wallets through a usage-pricing and transaction ledger, so cost is attributable per action and per tenant.
| METRIC | BEFORE | AFTER | CHANGE | WINDOW |
|---|---|---|---|---|
| AI capabilities in production | — | 10 across ~20 brands | Structural | 2025-08 to 2026-07 |
| Prompt, tool or model change | Code release | Configuration row update | Structural | Since 2026-04-23 |
| Agent tool calls with an audit record | — | 100%, with arguments, status and duration | Architectural | Since 2026-04-23 |
| Cross-tenant brand names in model output | — | Scanned and redacted on every turn | Architectural | Since 2026-04-23 |
| Advanced search | Rewards filter expertise | Plain-sentence query, same engine | Structural | Since October 2025 |
| Agents using an AI surface weekly | 35% — illustrative | 70% — illustrative | +35 percentage points — illustrative | Proposed comparison: 60 days either side of briefing cards shipping |
| Support tickets asking how a feature works | 40 per month — illustrative | 25 per month — illustrative | ~38% reduction — illustrative | Proposed comparison: three months before and after grounded help |
Two capabilities are worth reading on their own. The document parser is the only fully unattended pipeline in the platform, and the email approval queue is the strongest governance property in it.
05. Conclusion
AI shipped per tenant, not per release
Ten capabilities run on one provider, retrieval stays inside each tenant’s database, and reasoning capacity is spent where planning actually requires it. Staged per-tenant rollout is the release gate rather than percentage traffic, because the customers are agencies rather than anonymous users – one capability sat in live validation with a single agency for fourteen weeks before any tenant enablement. What remains is written down and sequenced, starting with the evaluation replay becoming a blocking gate.
Talk to Teamvoy About AI in Your Multi-Tenant Product
Tell us where AI output touches a customer in your system today