Semantic Layer for AI Agents in Multi-Tenant Analytics
Learn how a semantic layer for AI agents connects metrics, tenant policy, schema context, tools, and execution for reliable multi-tenant B2B SaaS analytics.
An analytics agent should not decide what revenue means, which tenant is active, and where SQL may run in the same model call. Those are separate contracts with separate owners.
Short answer: a semantic layer for AI agents is the contract between business meaning and executable analytics. For multi-tenant SaaS, that contract must connect six things: approved metrics, verified tenant policy, retrieved schema context, allowed tools, a controlled execution path, and evidence about what happened. A model can choose among those inputs. It should not be the authority that defines or enforces them.
This definition is broader than a metrics catalog and narrower than an entire analytics platform. It gives the agent enough governed meaning to act without giving it permission to invent joins, choose a tenant, or run arbitrary SQL.
Key takeaways
- A semantic layer for agents is an executable contract, not a larger prompt.
- Metric meaning and tenant policy are different. "Net revenue" defines a calculation.
tenantIddefines who may see the result. - MCP and other tool protocols describe how an agent calls a capability. They do not define metrics or make a tool safe.
- Schema retrieval should return the smallest useful slice: relevant tables, columns, glossary entries, and trusted query examples.
- The execution boundary matters as much as SQL generation. Headful and headless products can share semantics while executing in different places.
- Prompt-based tenant filtering needs downstream verification and database controls. A prompt alone is not an authorization system.
- Evaluate the contract with adversarial tenant, metric, tool, and failure cases, not only correct-answer demos.
The six-part semantic contract
An agent receives a question such as:
Compare activation rate for this customer with the previous quarter, then add the result to their dashboard.
That sentence hides six decisions. Treating them as one "text to SQL" task makes failures hard to detect.
| Contract | Question it answers | Authoritative source | What the agent may do |
|---|---|---|---|
| Metrics | What does "activation rate" mean? | Reviewed metric definition, glossary, or trusted SQL | Select the matching definition and explain assumptions |
| Tenant policy | Which customer's data may be read? | Server-verified identity, product permissions, database policy | Carry the supplied scope, never choose or widen it |
| Schema context | Which tables, columns, joins, and time fields are relevant? | Current schema plus curated annotations | Search, link, and prune context |
| Tools | Which actions exist, and what inputs do they accept? | Typed tool registry and authorization policy | Call only permitted tools with validated arguments |
| Execution | Where may SQL run, with which credentials and limits? | Query gateway, SDK callback, database role, and resource policy | Submit a bounded query, not bypass the executor |
| Evidence | How can a person reproduce or dispute the answer? | Logs, generated SQL, parameters, definition versions, result metadata | Return a rationale and trace identifiers appropriate to the audience |
The most common architecture mistake is collapsing these rows into one system prompt. The prompt then contains a metric definition, a tenant ID, schema text, safety instructions, and tool descriptions. That may improve generation, but it does not create an enforceable contract. Any rule that matters after the model responds needs a non-model owner.
Reference architecture for multi-tenant analytics agents
The safe path starts with identity, not language.
There are three trust transitions in this flow.
- Identity to policy: the SaaS backend converts an authenticated application session into a tenant and role. The browser does not get to make that mapping.
- Meaning to action: the agent maps the user's words to approved business and schema context, then proposes a tool call or query.
- Proposal to execution: deterministic code checks the proposal and runs it through a controlled database connection.
This separation also keeps UI decisions out of the data contract. A React dashboard, an MCP client, and a custom backend assistant may present different experiences while sharing metric definitions and tenant rules.
1. Metrics are versioned business logic
A metric contract should state more than a label and SQL fragment. For each important metric, record:
- definition and owner
- grain, such as workspace-day or invoice
- allowed dimensions
- time column and timezone
- default filters and exclusions
- valid join paths
- review date or version
- one or more known-good queries and test fixtures
For example, "active workspace" may mean at least one qualifying event in the last 28 complete days, excluding internal accounts. If an agent only sees workspaces.last_seen_at, it can produce valid SQL for the wrong definition.
Not every team needs a warehouse-scale semantic modeling product before testing an agent. A smaller contract can begin with glossary entries, table and column annotations, and gold SQL for the ten questions that matter most. The important property is reviewability. A human can inspect the definition and a test can prove the compiled query matches it.
2. Tenant policy is not another metric
Tenant policy answers a different question: which rows and actions belong to the authenticated caller?
The trusted input should come from a server-side session or a verified signed token. It may include:
- organization or workspace ID
- tenant ID
- user ID
- role or plan
- permitted datasources
- row or column restrictions
- allowed analytics actions
Do not retrieve tenant identity from a React prop, URL parameter, chat message, or model output. Those values can help the UI render state, but they are not authority.
The policy should survive every surface. If dashboards apply one tenant rule while an MCP tool and an AI assistant each reimplement another, the newest surface is likely to become the weakest one.
3. Schema context is retrieved, linked, and pruned
Sending an entire database schema wastes context and increases ambiguity. A better sequence is:
- retrieve likely table summaries
- retrieve columns for those candidate tables
- add related glossary entries and trusted SQL examples
- link phrases in the question to concrete schema elements
- remove unrelated context before generation
This turns "country" into an explicit mapping such as customers.ip_country instead of asking the model to choose among five plausible columns. It also gives the system a place to surface ambiguity before execution.
Context retrieval still needs tenant-aware metadata boundaries. Schema chunks for one organization must not appear in another organization's search results. Conversation history and agent memory need the same treatment. Data rows are not the only objects that can leak across tenants.
4. Tools expose narrow capabilities
An agent tool should represent one capability with typed input, a clear authorization check, and predictable errors.
Good analytics tools look like:
search_schema(question, datasource)list_metrics(datasource)generate_sql(question, context_ids)run_read_query(sql, params, tenant_context)create_chart(result_shape, chart_intent)save_dashboard_block(dashboard_id, block, tenant_context)
A single execute_anything(prompt) tool hides too much. The agent cannot tell discovery from execution, and reviewers cannot assign different permissions to read, query, and mutate operations.
MCP analytics can standardize tool discovery and invocation. It does not supply metric definitions, tenant authorization, query limits, or audit policy. Protocol and semantics solve different jobs.
5. Execution is a product boundary
Where SQL runs changes the security and data-flow model.
In a managed headful path, the analytics service may resolve datasource credentials, execute SQL, and return a finished chart to an embedded React workspace. That reduces integration work, but the service is part of the data path.
In a headless path, a cloud service can generate SQL while a server-side SDK validates and executes it inside the customer's backend. Credentials and raw query results can remain in that infrastructure. The customer then owns the UI, error states, persistence, and database limits.
Neither path is automatically safe. Both still need read-only credentials, database permissions, timeouts, row limits, parameter binding, and logs. "The agent generated it" is never a reason to bypass the normal query boundary.
What the semantic layer should compile
The agent should operate on a compact request that can be validated before SQL exists.
{
"metric": "activation_rate",
"dimensions": ["week"],
"timeRange": {
"preset": "previous_quarter"
},
"tenantContext": {
"source": "verified_jwt",
"tenantId": "tenant_from_server"
},
"output": {
"type": "line_chart"
}
}
The semantic and execution layers can then compile that request into parameterized SQL for the selected dialect. The tenant ID shown here is illustrative. It should be inserted by trusted server code, not copied from user text.
Free-form SQL generation can still be useful when the question does not fit a predefined metric API. In that case, treat SQL as a proposal. Verify table access, tenant predicates, parameters, statement type, complexity, and limits before execution.
Where QueryPanel fits, and where it does not
QueryPanel's primary product is its headful React SDK: a Notion-like dashboard workspace with an AI assistant for tenant-level customization. The customer backend mints an RS256 JWT with organization and tenant claims. The embedded client sends that token to the QueryPanel API, which verifies the claims and handles the managed query path.
This path is for teams that want customers to ask questions, create charts, and keep dashboard changes without building each interaction state. It is not a zero-trust execution path. In the headful architecture, QueryPanel's API resolves Vault-backed datasource credentials and executes queries for the embedded experience.
QueryPanel also offers a headless Node SDK for teams building a custom interface. The default v2 path sends the question and schema context to the QueryPanel API for SQL generation, then validates and executes the returned SQL through an adapter on the customer's server. Raw result rows and database credentials remain in customer infrastructure. Chart generation receives anonymized result shapes rather than raw values.
QueryPanel's current semantic context is practical and retrieval-based:
- schema-aware table and column chunks
- organization-scoped hybrid vector and full-text retrieval
- glossary entries
- gold SQL examples
- entity-to-schema linking and context pruning
- SQL reflection before execution
It is not a LookML, MetricFlow, or Cube-style metrics engine. If your organization already has a governed semantic model, keep it as the metric authority and connect the agent to it. If you do not, QueryPanel's glossary, annotations, and gold SQL provide a smaller starting point for grounded customer-facing analytics.
Tenant filtering also needs a precise description. QueryPanel instructs the SQL model to bind a configured tenant field, then verifies the generated SQL and parameters. The headless SDK has an additional execution safeguard that can add a missing tenant predicate. This is defense in depth around a prompt-based generation step, not a replacement for database grants, row-level security, read-only roles, or tenant-isolated test data.
For the full integration sequence, see how to add an AI analytics assistant to a React SaaS app. For broader SQL safety and cost controls, use the production NL-to-SQL checklist.
Seven failure modes to test before launch
| Failure mode | What breaks | Test case | Required control |
|---|---|---|---|
| Metric drift | Two surfaces calculate the same KPI differently | Compare dashboard, agent, and approved SQL for one fixture | Versioned definition and contract tests |
| Tenant chosen by the client | A modified prop or request changes account scope | Replace the browser tenant value | Server-derived identity and signed context |
| Prompt-only isolation | Generated SQL omits or rewrites the tenant filter | Ask a broad question with no tenant wording | Post-generation verification plus DB policy |
| Context collision | Similar columns from unrelated tables are selected | Seed two plausible country and revenue fields | Retrieval filters, schema linking, ambiguity handling |
| Tool overreach | A read question invokes a write or broad execution tool | Ask the agent to "fix" a number | Narrow tools and per-tool authorization |
| Execution bypass | Generated SQL avoids normal limits or roles | Produce a large scan or forbidden table reference | One query gateway with limits and allowlists |
| Cross-tenant memory | A follow-up recalls another customer's context | Reuse or swap session identifiers across tenants | Tenant-scoped sessions, vectors, caches, and logs |
Two seeded tenants make isolation failures visible. Give Tenant A exactly 11 qualifying records and Tenant B exactly 29. Run the same broad prompt for both, then try a tenant switch from the browser and a reused conversation ID. The expected result is obvious without reading hundreds of rows.
Evaluation checklist for a semantic layer for AI agents
Score the system end to end. A correct query in a notebook is not enough.
Meaning
- Each launch metric has an owner, definition, grain, time rule, and exclusions.
- The agent can distinguish two plausible definitions or ask for clarification.
- Approved questions have gold SQL or another deterministic expected result.
- Metric changes have versions and regression tests.
Tenant policy
- Tenant identity comes from a server session or verified token.
- Dashboard, agent, MCP, cache, memory, and query paths use the same scope.
- Broad questions remain tenant-filtered.
- Database permissions limit damage if application checks fail.
Context
- Retrieval is filtered by organization and datasource.
- The agent sees a small relevant schema slice, not the full catalog.
- Business terms link to explicit tables and columns.
- Schema changes trigger sync and stale-context tests.
Tools and execution
- Discovery, generation, execution, and mutation are separate capabilities.
- Every tool validates inputs and checks authorization.
- Only read statements reach the analytics executor.
- Queries use parameters, read-only roles, timeouts, and row limits.
- Errors fail closed without exposing credentials or internal schema to customers.
Evidence and user experience
- Logs connect the user question, tenant, context IDs, SQL, parameters, and result metadata.
- Customer rationales use business language without exposing tenant IDs or database paths.
- Admins can inspect enough detail to reproduce a disputed answer.
- Empty, ambiguous, denied, and failed requests have deliberate UI states.
- A useful answer can become a saved chart or dashboard view when the product promises that workflow.
Cube's September 3, 2026 agentic analytics harness article makes a useful distinction between an executable semantic layer and the surrounding context an agent needs. That warehouse-centered pattern is a strong fit when certified measures, dimensions, joins, and access policies already live in a semantic model. The checklist above adds the SaaS product contracts that still sit around it: tenant identity, UI behavior, tool permissions, and the location of execution.
FAQ
What is a semantic layer for AI agents?
A semantic layer for AI agents is an executable contract that connects business definitions to allowed analytics actions. It tells an agent which metrics, dimensions, joins, tenant rules, and tools are valid, then gives deterministic code enough information to compile and verify the request.
Is a semantic layer the same as a vector database?
No. Vector search can retrieve relevant schema descriptions, glossary entries, and examples. It does not enforce a metric calculation or access policy by itself. Retrieval is one source of context inside the contract.
Does MCP provide a semantic layer?
No. MCP standardizes how an agent discovers and calls tools. A semantic layer defines business meaning and often compiles governed requests. An MCP server can expose a semantic layer, but the protocol does not create one.
How should tenant isolation work for an analytics agent?
Derive tenant identity on the server, carry it through a verified token or trusted backend call, apply it during query generation, verify it before execution, and enforce database permissions underneath. Scope sessions, caches, vector retrieval, and agent memory by tenant too.
Is prompt-based tenant filtering safe?
Not by itself. Prompt instructions can help generate the right predicate, but prompts are probabilistic. Add deterministic SQL and parameter checks, a controlled execution gateway, read-only database roles, and row-level policy where the database supports it.
Do we need a full metrics engine before using an analytics agent?
No. Start with reviewed definitions, annotations, glossary terms, and gold SQL for a small set of customer questions. Use a full metrics engine when the organization needs one governed model across many tools, teams, and data products.
When should we choose a headful or headless QueryPanel path?
Choose the headful React SDK when you want a complete embedded dashboard workspace, AI assistant, and customer customization with less UI work. Choose the headless Node SDK when you need a custom interface and want SQL execution and raw results to stay in your backend.
Build the semantic contract before widening the agent's permissions. Start with QueryPanel's headful React experience for a complete customer analytics workspace, or use the headless Node SDK when your product needs a custom interface and local execution boundaries.