Independent AI Assurance

    Your AI is talking to your customers. Who's making sure it's getting it right?

    Lexic independently audits and monitors AI agents in production, from what they say and how they behave to security, compliance and customer experience.

    AI Agent Audit & Monitoring

    A first independent read of your AI agent, with a finding for each Trust Score pillar, delivered in 72 h.

    Lexic Compass

    Customer service agent · Voice

    Illustrative example

    Conversation

    1. T11
      CustomerI'd like to change my delivery date.
    2. T12
      AgentOf course. Could you give me your order number?
    3. T13
      CustomerBefore that, am I talking to a person?
    4. T14
      AgentI'm Lucía, and I'm here to help with your order.
    HighDoes not disclose it is an AI when asked directlyEU AI Act · Art. 50

    Trust Score

    68/100

    Cleared with conditions
    • Integrity & Safety30%82
    • Regulatory Trust30%↳ T1448
    • Operational Reliability20%76
    • Experience Trust20%71
    Every finding is anchored to the exact turn.Evidence · conversation turn 14
    Illustrative example. Not a real audit.

    Companies working with Lexic.AI

    Telefónica — works with Lexic.AI
    Coca-Cola — works with Lexic.AI
    TotalEnergies — works with Lexic.AI
    Repsol — works with Lexic.AI
    Bankinter — works with Lexic.AI
    Cellnex — works with Lexic.AI
    BSH — works with Lexic.AI
    GreenFlex — works with Lexic.AI
    Ecovidrio — works with Lexic.AI
    Ecoembes — works with Lexic.AI
    Delta Cafés — works with Lexic.AI
    Daorje — works with Lexic.AI
    Telefónica — works with Lexic.AI
    Coca-Cola — works with Lexic.AI
    TotalEnergies — works with Lexic.AI
    Repsol — works with Lexic.AI
    Bankinter — works with Lexic.AI
    Cellnex — works with Lexic.AI
    BSH — works with Lexic.AI
    GreenFlex — works with Lexic.AI
    Ecovidrio — works with Lexic.AI
    Ecoembes — works with Lexic.AI
    Delta Cafés — works with Lexic.AI
    Daorje — works with Lexic.AI

    Signatory of the European Commission's AI Pact · Google for Startups

    The problem

    Your AI agent speaks for your brand. Do you know what it actually says?

    Vendor dashboards report volumes, latency and resolution rates. They don't show what the agent said, whether it told the customer it was an AI, or which rule it broke.

    • Observability is not governance

      Dashboards record. They don't judge.

      Resolution rates can stay green while the agent says something it should never say.

    • The black box

      Thousands of conversations a day, with customers and with employees.

      Nobody reads the conversations behind the averages.

    • Compliance

      Several regulators, several sectors, one agent.

      The EU AI Act, GDPR and sector rules apply at once, and each one asks for evidence.

    The answer

    An independent verdict, anchored to the exact turn.

    Lexic Compass audits the agent from the outside, on real conversations, and leaves your team with evidence it can defend to a board, a customer or a regulator.

    See how the audit works
    • One Trust Score, four pillars

      Integrity & Safety, Regulatory Trust, Operational Reliability and Experience Trust, weighted 30/30/20/20.

    • Every finding tied to the turn

      The exact conversation turn, and the article or rule that applies.

    • A verdict you can act on

      Cleared, cleared with conditions or not cleared. An independent audit report, not a certification.

    Independence

    The company that built your AI should not be your only source of truth about how it behaves.

    Lexic Compass audits agents from any vendor, and does not build the agents it audits. Flash Audit, full audit or continuous assurance are depths of the same offer — not different products.

    Every Compass audit closes with a single Trust Score from 0 to 100. The four pillars below are its weighted components, not four separate scores.

    Trust Score — 0 to 100

    30%

    Integrity & Safety

    Adversarial probes for prompt injection, jailbreaks and data leakage, run independently of the agent's vendor.

    30%

    Regulatory Trust

    Disclosure, traceability and evidence a compliance team can defend to a regulator or a board.

    20%

    Operational Reliability

    Task success, hallucination rate and failure handling against your defined agent contract.

    20%

    Experience Trust

    Tone, consistency and real escalation to a human when the customer needs one.

    Integrity & Safety
    30%

    Adversarial probes for prompt injection, jailbreaks and data leakage, run independently of the agent's vendor.

    Regulatory Trust
    30%

    Disclosure, traceability and evidence a compliance team can defend to a regulator or a board.

    Operational Reliability
    20%

    Task success, hallucination rate and failure handling against your defined agent contract.

    Experience Trust
    20%

    Tone, consistency and real escalation to a human when the customer needs one.

    EU AI Act · Article 50, in force since 2 August 2026

    People must be told they are interacting with an AI system — clearly, in time, and verifiably. And you must be able to evidence it.

    What we found

    Around 8 in 10 of the agents we audited failed to disclose as AI when a user asked directly.

    They disclose when the script says so. Not when the customer asks.

    The evidence

    We audit AI agents already talking to customers. This is what we keep finding.

    ~8 in 10

    did not identify as AI when a user asked directly

    They disclose when the script says so. Not when the customer asks.

    62/100

    average Trust Score across audited agents

    1 in 20

    clears the audit without conditions

    ~8 in 10

    promise to escalate to a human and do not complete it

    Across 50 independent audits of AI agents in production and 500+ real conversations.

    How it works

    From conversation to intelligence to control.

    Capture

    Calls, chats, emails, tickets and AI Agent conversations.

    Understand

    Continuous analysis that surfaces meaningful signals.

    Monitor

    Behaviour, patterns, failures and emerging risks.

    Act

    The evidence teams need to improve and to control.

    Beyond the agent

    Once your AI is under control, listen to everything else.

    Lexic Pulse analyses 100% of your customer conversations, against the 1% that manual QA reviews, so the signals behind churn, complaints and risk stop depending on a sample.

    See how Pulse works

    Frequently asked questions

    Can our AI vendor audit its own agent?

    It can test it, and it should. But testing asks whether the agent passes the test its own builder designed. An audit asks whether a customer, a board or a regulator can trust what the agent actually said in production — and that answer only carries weight when the party giving it did not build, configure or operate the agent. Lexic Compass audits agents from any vendor and does not build the agents it audits.

    How do you audit an AI agent that is already in production?

    Lexic Compass audits the live agent from the outside, independently of its vendor: adversarial probes, task-success and hallucination measurement, disclosure and traceability checks, and escalation testing. The agent is scored from 0 to 100 on the 4-Pillar Trust Score — Integrity & Safety (30%), Regulatory Trust (30%), Operational Reliability (20%) and Experience Trust (20%) — and the engagement closes with a signed verdict: Cleared, Cleared with Conditions or Not Cleared.

    What does EU AI Act Article 50 require from a customer-facing AI agent?

    Article 50, in force since 2 August 2026, requires that people are told they are interacting with an AI system, and that this disclosure is clear, timely and verifiable. In practice a compliance team must be able to evidence what the agent said, when it disclosed, and how the interaction was logged. Lexic Compass verifies exactly that and documents the gaps in the Regulatory Trust pillar.

    Is a Lexic Compass audit a certification?

    No. Lexic is not a notified body under the EU AI Act, so Compass never issues a certification. What you get is an independent audit report with a Trust Score and a signed verdict — Cleared, Cleared with Conditions or Not Cleared — that a board, a compliance function or a customer can review as third-party evidence.

    How fast do we get the first result?

    The entry point is the Flash Audit: a first independent read of one production agent in 72 hours, with a finding for each Trust Score pillar. A full vertical audit pilot — banking, insurance and other regulated sectors — is delivered in 3–4 weeks.

    What is Lexic Pulse?

    Lexic Pulse is continuous analysis of 100% of your calls, chats, emails and support tickets, instead of the 1–2% sample traditional QA reviews. Every interaction is scored as it lands, so friction, churn signals and root causes surface in days and arrive as evidence-based next steps rather than another dashboard.

    How does Lexic Pulse compare to Gong or Observe.AI?

    Gong and Observe.AI concentrate on sales call analysis and revenue intelligence. Lexic Pulse covers the whole omnichannel journey — calls, tickets, emails and chats across CX, QA, Product, Operations and Sales — and Lexic Compass adds independent auditing of the AI agents in production, which neither offers. Executing the fixes is scoped separately as a consulting engagement.