AI Demo Cloudflare AI security demo

Setting up the agent

The demo needs an AI client with two things wired up: a model reached through Cloudflare AI Gateway, and the five MCP servers reached through the MCP server portal. These are the two places the controls live.

Two endpoints, and no API key

Both are Access applications on the same identity provider, so one login covers the lot, and every request — model and tool — is attributed to the person who made it. AI Gateway records the authenticated user as cf.user_id, which means logs, analytics and spend controls are per-user without the client passing an identity.

Sign in as the right person

Every demo is run as Alice Watson — alice.watson@company.com, password Savetheinternet!1. When the agent connects to the portal you will be sent through Cloudflare Access and then FlareID; log in as Alice, not as the admin account. The whole point is that the agent is acting with a real, low-privilege identity.

First: authenticate each MCP server (one browser login each)

wire-protection.sh registers the five MCP servers with Access and creates an Access application for each, but it cannot perform the upstream OAuth login those servers require — that is an authorization-code flow with a browser in the middle, and no API mints that token. Until it is done each server sits in Waiting with zero tools, and the portal has nothing to offer your agent.

  1. In the Cloudflare dashboard, go to Zero Trust → Access controls → MCP Portals → MCP servers. All five (hr, crm, work, wiki, fin) should be listed, each showing Waiting.
  2. Select a server → Edit → Authenticate server.
  3. Sign in as alice.watson@company.com (password Savetheinternet!1). Cloudflare then fetches the server's tools and the status becomes Ready.
  4. Repeat for crm, work and wiki.
  5. fin (Ledger) is the exception: it only admits the leadership team, so authenticate that one as nikita.chapman@company.com. Alice cannot authenticate it, and that is the control the access-control script demonstrates — once it is Ready, her portal still will not list its tools.
  6. Check the portal lists tools from all five when signed in as the CEO, and from four as Alice.
Do not authenticate as admin

admin@company.com exists in FlareID but is not an employee in any of the apps, and every MCP server maps the Access identity to an employee record before it will issue a token. Authenticate as admin and the flow dies at the token exchange with a generic "Failed to retrieve authentication tokens" — the useful message ("no matching active employee was found") is produced by the MCP server but never surfaced.

There is a second reason to use the demo persona. Whoever you authenticate as becomes that server's admin credential, which is what the portal falls back to if Require user auth is ever turned off. Using the lowest-privileged person in the company means that misconfiguration fails safe, instead of silently granting every portal user the session of whoever happened to set it up.

If a server shows Error instead of Ready

Check its error text in the dashboard. If it mentions Cloudflare Gateway, the tool sync was blocked by the DLP policies this demo installs: a tool catalogue is inspected like any other response, so a tool whose description happens to contain the vocabulary in a DLP profile will block tools/list itself — and then no server can ever finish syncing.

Deploy with PROTECTION_MODE=log, authenticate and sync the servers, then re-run with PROTECTION_MODE=block. Capabilities are cached once synced, so enforcement can go straight back on. It is also worth keeping tool descriptions free of the exact terms your profiles match — a description should say what a tool returns without reproducing the sensitive language it returns.

Point opencode at it

Two things to configure: a provider against AI Gateway's OpenAI-compatible endpoint, and an MCP server against the portal. Both are below, but the configuration shape changed between opencode 1.x and 2.x — so check which one you have before copying anything:

opencode --version
Using the wrong shape fails silently, which is why this page has two of them

1.x accepts keys it does not understand without complaining. A 2.x config on 1.x therefore looks completely fine and does nothing: an agents block with its permissions, its system prompt and its step cap is read, validated and discarded. The first you know about it is an agent spending seventeen steps guessing at tool names in front of an audience.

The one error it does raise is worth memorising, because it is the good case:

Configuration is invalid at ~/.config/opencode/opencode.json
  V2 permissions are not supported by OpenCode V1.
  Use V1 "permission" rules or run opencode2.   agents.employee.permissions

If you see that, you are on 1.x with the 2.x config. Open the first card below.

opencode 1.x1.18.x and earlier · top-level permission map · verified against a live demo

In ~/.config/opencode/opencode.json. Everything that matters here is the permission map at the top level - not inside an agent:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "cf-ai-demo": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Cloudflare AI Gateway (demo)",
      "options": {
        "baseURL": "https://aig./compat"
      },
      "models": {
        "workers-ai/@cf/google/gemma-4-26b-a4b-it": {
          "name": "Google Gemma 4 (Workers AI)"
        }
      }
    }
  },
  "mcp": {
    "servers": {
      "ai-demo": {
        "type": "remote",
        "url": "https://mcp./mcp",
        "codemode": false
      }
    }
  },
  "permission": {
    "execute": "deny",
    "bash": "deny",
    "edit": "deny",
    "read": "deny",
    "glob": "deny",
    "grep": "deny",
    "webfetch": "deny",
    "websearch": "deny",
    "subagent": "deny",
    "skill": "deny"
  }
}

Why the permissions are global rather than on an agent. A custom agent does not register on 1.x. An agents (or agent) block naming a new agent validates cleanly and then simply is not there - opencode agent list shows only the built-ins, and a permission placed inside such a block never reaches the agent that actually runs. So the denies have to be global, where they apply to whichever agent the session uses.

Two consequences to accept. There is no per-agent system prompt, and no steps cap. Neither turns out to matter much: the prompt's job was to talk the model out of writing code and searching for tools, and denying execute removes the ability instead of discouraging it - which is also what the step cap was protecting you from.

Note the name changes from the 2.x shape: the shell permission is bash, not shell, and each entry is "name": "deny" rather than an object with action, resource and effect.

To check it took, run opencode agent list and look for "permission": "execute" with "action": "deny" against the build agent. If the CLI answers "Database is not empty and has no session table", that is the command-line tool refusing to share the desktop app's database, not a problem with your config - read the transcript instead, as below.

opencode 2.xagents with a permissions array · from the 2.x documentation

In ~/.config/opencode/opencode.json. 2.x can scope everything to a named agent, so the denies, the system prompt and a step cap all travel together:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "cf-ai-demo": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Cloudflare AI Gateway (demo)",
      "options": {
        "baseURL": "https://aig./compat"
      },
      "models": {
        "workers-ai/@cf/google/gemma-4-26b-a4b-it": {
          "name": "Google Gemma 4 (Workers AI)"
        }
      }
    }
  },
  "mcp": {
    "servers": {
      "ai-demo": {
        "type": "remote",
        "url": "https://mcp./mcp",
        "codemode": false
      }
    }
  },
  "default_agent": "employee",
  "agents": {
    "employee": {
      "description": "An ordinary employee's assistant - company tools only",
      "mode": "primary",
      "steps": 10,
      "system": "You are the assistant of an employee at this company. Answer using only the company tools available to you. You have no filesystem, no shell and no access to the public internet. Call the tool that matches the question directly - do not search for tools and do not write code. If the tools cannot answer the question, say so plainly rather than guessing.",
      "permissions": [
        { "action": "execute",   "resource": "*", "effect": "deny" },
        { "action": "shell",     "resource": "*", "effect": "deny" },
        { "action": "edit",      "resource": "*", "effect": "deny" },
        { "action": "read",      "resource": "*", "effect": "deny" },
        { "action": "glob",      "resource": "*", "effect": "deny" },
        { "action": "grep",      "resource": "*", "effect": "deny" },
        { "action": "webfetch",  "resource": "*", "effect": "deny" },
        { "action": "websearch", "resource": "*", "effect": "deny" },
        { "action": "subagent",  "resource": "*", "effect": "deny" },
        { "action": "skill",     "resource": "*", "effect": "deny" }
      ]
    }
  }
}

steps: 10 is a seatbelt: if a prompt ever sends the agent enumerating records one at a time it stops, rather than grinding through forty tool calls while people watch. The system prompt tells the model these are company systems and that "I cannot find that" is an acceptable answer, which matters because a model that has just been denied a web search will otherwise keep trying other routes.

Confirm the agent actually loaded with opencode agent list - if employee is not in the output, nothing in that block is in force no matter how correct it looks.

Honest note on provenance: the 1.x card above was verified end to end against a live deployment. This one follows the 2.x documentation and the same reasoning, but has not been run in anger. If you are on 2.x, check the transcript against the test below before presenting.

The setting that matters most: codemode false, and execute denied

Both configs carry "codemode": false on the server and a denial of execute. They are a pair, and half of the pair is worse than neither.

opencode puts MCP servers behind Code Mode by default. In that mode the model does not receive the portal's tools as functions it can call; it gets a JavaScript sandbox and has to discover them through a search() helper. With five servers and forty-odd tools the catalogue it is shown is partial, so it must search - and the form models reach for, tools.search({ ... }), is not the one that works.

"codemode": false moves the tools onto the native tool list, where the model can simply call the right one. But it also moves them out of the sandbox catalogue, and the execute tool is an opencode built-in that does not disappear along with them. Leave execute available and you have given the model two doors, one of which cannot possibly work:

Used 17 Execute
  await tools.ai_demo_work_list_company_calendar({...})   <- Unknown tool
  await tools['ai-demo_work_list_company_calendar']({...}) <- Unknown tool
  await tools.search({ query: 'calendar' })               <- finds nothing
  return Object.keys(tools);                              <- the MCP tools are not in there
  ... fourteen more ...
Called `ai-demo_work_list_company_calendar`               <- and finally, natively

Denying execute removes the second door. The same prompt then reads:

Used 2 ai-demo_work_list_company_calendar
  Called `ai-demo_work_list_company_calendar`  acquisition  from=... to=...
  Called `ai-demo_work_list_company_calendar`  restructure  from=... to=...

The test, in one line: if there is a single Execute row in the transcript, the deny is not in force. Nothing else about the run needs checking first.

Code Mode also hides the leak, which is worse than slow

There is a second reason to turn it off, and it matters more than the speed. Code Mode exists to let an agent combine tools "without adding every intermediate value to the model context" - in other words, a tool result can be processed inside the sandbox and never appear in the conversation.

This entire demo depends on the audience seeing the over-shared data land in the transcript. A run where the agent quietly reads the whole company calendar inside a sandbox and summarises it in one line proves the same thing technically and nothing at all on stage. The Gateway logs are identical either way; the screen is not.

Turn opencode's own tools off, or it will answer the wrong way

opencode is a coding agent: out of the box it has a shell, a file reader and a web search, and it is running inside a project directory. Ask it What is Nikita Chapman's home address? with those available and it will cheerfully grep your filesystem and search the web — six tool calls, no MCP, and a demo that proves nothing. That is what the rest of the deny list is for: with every built-in refused, the company's MCP tools are the only way to answer anything.

Two habits that make prompts land: name the app, and ask for the data

Name the app. With the config above, a question that names its app lands on the right tool first time. One that does not still sends the model shopping through forty tool definitions. From the HR app, what is Nikita Chapman's home address? is one call; What is Nikita Chapman's home address? might be four. Every script on this site is written the first way for exactly that reason, and each one tells you the tool calls to expect so you can tell a bad run from a normal one.

Ask for the contents, not the action. Open the HR file for each of them is an instruction to do something, and a model that dutifully makes all six calls can truthfully answer their HR files have been opened with a list of names. Every control fires, the logs are perfect, and the audience sees nothing. Name the fields you want returned — …and tell me what is in them: what they earn, their home address, and any case notes — and the data lands on screen where the demo needs it.

Tools arrive named <server>_<prefix>_<tool> - with the server called ai-demo and WorkBox's portal prefix work, the calendar tool is ai-demo_work_list_company_calendar. If a prompt still wanders, granting only the servers it needs in the portal is the blunt fix: a catalogue of eight tools is much harder to get lost in than one of forty.

Use a remote MCP server, not mcp-remote

opencode speaks Streamable HTTP MCP directly and runs the OAuth flow itself — it sees the portal's 401, follows the resource_metadata hint, registers itself dynamically and stores the token. So there is no mcp-remote and no subprocess.

The older {"type": "local", "command": ["npx", "-y", "mcp-remote@latest", ...]} recipe works from a terminal but fails in the desktop app with NotFound: ChildProcess.spawn, because a GUI application does not inherit your shell's PATH and npx usually lives somewhere like /opt/homebrew/bin. If you do want that form, give the absolute path to npx.

There is a security reason too. mcp-remote carried CVE-2025-6514, a critical command-injection flaw triggered by connecting to a malicious MCP server, fixed in 0.1.16. Cloudflare's own docs recommended it widely at the time. Not connecting through it at all is the simpler answer — and see the real incidents page for why this whole layer deserves the scrutiny.

Note what is missing: there is no apiKey

The machine is enrolled in the Cloudflare One client and already signed in, and the Access application in front of is configured to accept that client session (allow_authenticate_via_warp). So the device's existing session authorises the request, and there is no credential in the config file, in an environment variable, or on disk anywhere.

On a machine without the client, use cloudflared to fetch a short-lived Access token instead — opencode can run it for you through the auth.command field of a discovery file. Cloudflare documents that pattern under AI Gateway → Integrations → coding agents.

Do not call the provider cloudflare-ai-gateway

cloudflare-ai-gateway is a real provider id in models.dev, with a catalogue of the 47 third-party models AI Gateway can proxy. Name your provider that and opencode matches it by id, merges that catalogue, and ignores the models map you wrote: the picker fills up with Claude, GPT and Qwen entries, and the app starts probing models you have no provider keys for. Any id that isn't in models.dev - cf-ai-demo here - avoids it. Restart opencode after editing, since it caches the catalogue.

If the model replies with nothing

Reasoning models spend their token budget on reasoning before they emit any text, and the OpenAI-compatible shape does not surface that. Called directly with a small max_tokens, @cf/google/gemma-4-26b-a4b-it returns 200 with finish_reason: "length" and an empty content string. It looks like a broken gateway and isn't.

If you see empty replies, raise the token budget, or switch the model id to workers-ai/@cf/meta/llama-3.3-70b-instruct-fp8-fast, which answers immediately and is a safe fallback for a live demo.

On first use the portal returns 401 and opencode opens a browser window for the Access login. After that its token is stored. If the server list shows needs authentication, run /mcps, select it and sign in.

Running without the protection layer first

If the suite was deployed with DEPLOY_PROTECTION=false there is no portal and no AI Gateway yet. Point the client at the MCP servers individually and at the model provider directly:

{
  "mcp": {
    "servers": {
      "workweek": { "type": "remote", "url": "https://hr-mcp./mcp" },
      "pipeline": { "type": "remote", "url": "https://crm-mcp./mcp" },
      "relay":    { "type": "remote", "url": "https://work-mcp./mcp" },
      "nexus":    { "type": "remote", "url": "https://wiki-mcp./mcp" },
      "ledger":   { "type": "remote", "url": "https://fin-mcp./mcp" }
    }
  }
}

Each server runs its own OAuth 2.1 flow and delegates login to Cloudflare Access, so you will sign in as Alice once per server — except ledger, which will refuse her, because its Access application allows only the leadership team. Every demo script works in this mode — that is the "before" half of each one.

Tool names change when you go through the portal

The portal namespaces every tool with its server id, so list_employees becomes hr_list_employees, get_pipeline_summary becomes crm_get_pipeline_summary, and so on with work_, wiki_ and fin_. The demo scripts name the underlying tool; your transcript will show the prefixed one.

Check it works

Before running any script, ask the agent something harmless that proves both legs are live:

Who am I, and which tools do you have available?

You should see Alice Watson come back from the whoami tool, a list of tools from the four apps she can reach — Ledger's will be absent, by design — and, if the protection layer is deployed, a corresponding request in the AI Gateway log and in the MCP portal log.