Setting up the agent
The demo needs an AI client with two things wired up: a model reached through Cloudflare AI Gateway, and the five MCP servers reached through the MCP server portal. These are the two places the controls live.
- Model traffic goes to
, the AI Gateway's own custom domain, whichwire-protection.shputs behind Cloudflare Access. On a custom domain, AI Gateway accepts a valid Access JWT as the request credential, so the client sends no gateway token and no API key at all. - Tool traffic goes to the MCP portal at
, which fronts,,and.
Both are Access applications on the same identity provider, so one login covers the lot, and
every request — model and tool — is attributed to the person who made it. AI Gateway
records the authenticated user as cf.user_id, which means logs, analytics and spend
controls are per-user without the client passing an identity.
Sign in as the right person
Every demo is run as Alice Watson — alice.watson@company.com,
password Savetheinternet!1. When the agent connects to the portal you will be sent
through Cloudflare Access and then FlareID; log in as Alice, not as the admin account. The whole point
is that the agent is acting with a real, low-privilege identity.
First: authenticate each MCP server (one browser login each)
wire-protection.sh registers the five MCP servers with Access and creates an Access
application for each, but it cannot perform the upstream OAuth login those servers
require — that is an authorization-code flow with a browser in the middle, and no API mints
that token. Until it is done each server sits in Waiting with zero tools, and the
portal has nothing to offer your agent.
- In the Cloudflare dashboard, go to Zero Trust → Access controls → MCP Portals
→ MCP servers. All five (
hr,crm,work,wiki,fin) should be listed, each showing Waiting. - Select a server → Edit → Authenticate server.
- Sign in as
alice.watson@company.com(passwordSavetheinternet!1). Cloudflare then fetches the server's tools and the status becomes Ready. - Repeat for
crm,workandwiki. fin(Ledger) is the exception: it only admits the leadership team, so authenticate that one asnikita.chapman@company.com. Alice cannot authenticate it, and that is the control the access-control script demonstrates — once it is Ready, her portal still will not list its tools.- Check the portal lists tools from all five when signed in as the CEO, and from four as Alice.
admin@company.com exists in FlareID but is not an employee in any of the
apps, and every MCP server maps the Access identity to an employee record before it will
issue a token. Authenticate as admin and the flow dies at the token exchange with a generic
"Failed to retrieve authentication tokens" — the useful message
("no matching active employee was found") is produced by the MCP server but never surfaced.
There is a second reason to use the demo persona. Whoever you authenticate as becomes that server's admin credential, which is what the portal falls back to if Require user auth is ever turned off. Using the lowest-privileged person in the company means that misconfiguration fails safe, instead of silently granting every portal user the session of whoever happened to set it up.
Check its error text in the dashboard. If it mentions Cloudflare Gateway, the tool sync was
blocked by the DLP policies this demo installs: a tool catalogue is inspected like any other
response, so a tool whose description happens to contain the vocabulary in a DLP profile
will block tools/list itself — and then no server can ever finish syncing.
Deploy with PROTECTION_MODE=log, authenticate and sync the servers, then re-run
with PROTECTION_MODE=block. Capabilities are cached once synced, so enforcement can
go straight back on. It is also worth keeping tool descriptions free of the exact terms your
profiles match — a description should say what a tool returns without reproducing the
sensitive language it returns.
Point opencode at it
Two things to configure: a provider against AI Gateway's OpenAI-compatible endpoint, and an MCP server against the portal. Both are below, but the configuration shape changed between opencode 1.x and 2.x — so check which one you have before copying anything:
opencode --version
1.x accepts keys it does not understand without complaining. A 2.x config on 1.x therefore looks
completely fine and does nothing: an agents block with its permissions, its system
prompt and its step cap is read, validated and discarded. The first you know about it is an agent
spending seventeen steps guessing at tool names in front of an audience.
The one error it does raise is worth memorising, because it is the good case:
Configuration is invalid at ~/.config/opencode/opencode.json
V2 permissions are not supported by OpenCode V1.
Use V1 "permission" rules or run opencode2. agents.employee.permissions
If you see that, you are on 1.x with the 2.x config. Open the first card below.
opencode 1.x1.18.x and earlier · top-level permission map · verified against a live demo
In ~/.config/opencode/opencode.json. Everything that matters here is the
permission map at the top level - not inside an agent:
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"cf-ai-demo": {
"npm": "@ai-sdk/openai-compatible",
"name": "Cloudflare AI Gateway (demo)",
"options": {
"baseURL": "https://aig./compat"
},
"models": {
"workers-ai/@cf/google/gemma-4-26b-a4b-it": {
"name": "Google Gemma 4 (Workers AI)"
}
}
}
},
"mcp": {
"servers": {
"ai-demo": {
"type": "remote",
"url": "https://mcp./mcp",
"codemode": false
}
}
},
"permission": {
"execute": "deny",
"bash": "deny",
"edit": "deny",
"read": "deny",
"glob": "deny",
"grep": "deny",
"webfetch": "deny",
"websearch": "deny",
"subagent": "deny",
"skill": "deny"
}
}
Why the permissions are global rather than on an agent. A custom agent does not
register on 1.x. An agents (or agent) block naming a new agent validates
cleanly and then simply is not there - opencode agent list shows only the built-ins,
and a permission placed inside such a block never reaches the agent that actually runs. So the denies
have to be global, where they apply to whichever agent the session uses.
Two consequences to accept. There is no per-agent system prompt, and no
steps cap. Neither turns out to matter much: the prompt's job was to talk the model out
of writing code and searching for tools, and denying execute removes the ability
instead of discouraging it - which is also what the step cap was protecting you from.
Note the name changes from the 2.x shape: the shell permission is
bash, not shell, and each entry is
"name": "deny" rather than an object with action,
resource and effect.
To check it took, run opencode agent list and look for
"permission": "execute" with "action": "deny" against the
build agent. If the CLI answers
"Database is not empty and has no session table", that is the command-line tool refusing to
share the desktop app's database, not a problem with your config - read the transcript instead, as
below.
opencode 2.xagents with a permissions array · from the 2.x documentation
In ~/.config/opencode/opencode.json. 2.x can scope everything to a named agent, so the
denies, the system prompt and a step cap all travel together:
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"cf-ai-demo": {
"npm": "@ai-sdk/openai-compatible",
"name": "Cloudflare AI Gateway (demo)",
"options": {
"baseURL": "https://aig./compat"
},
"models": {
"workers-ai/@cf/google/gemma-4-26b-a4b-it": {
"name": "Google Gemma 4 (Workers AI)"
}
}
}
},
"mcp": {
"servers": {
"ai-demo": {
"type": "remote",
"url": "https://mcp./mcp",
"codemode": false
}
}
},
"default_agent": "employee",
"agents": {
"employee": {
"description": "An ordinary employee's assistant - company tools only",
"mode": "primary",
"steps": 10,
"system": "You are the assistant of an employee at this company. Answer using only the company tools available to you. You have no filesystem, no shell and no access to the public internet. Call the tool that matches the question directly - do not search for tools and do not write code. If the tools cannot answer the question, say so plainly rather than guessing.",
"permissions": [
{ "action": "execute", "resource": "*", "effect": "deny" },
{ "action": "shell", "resource": "*", "effect": "deny" },
{ "action": "edit", "resource": "*", "effect": "deny" },
{ "action": "read", "resource": "*", "effect": "deny" },
{ "action": "glob", "resource": "*", "effect": "deny" },
{ "action": "grep", "resource": "*", "effect": "deny" },
{ "action": "webfetch", "resource": "*", "effect": "deny" },
{ "action": "websearch", "resource": "*", "effect": "deny" },
{ "action": "subagent", "resource": "*", "effect": "deny" },
{ "action": "skill", "resource": "*", "effect": "deny" }
]
}
}
}
steps: 10 is a seatbelt: if a prompt ever sends the agent enumerating records one at
a time it stops, rather than grinding through forty tool calls while people watch. The
system prompt tells the model these are company systems and that "I cannot find that" is
an acceptable answer, which matters because a model that has just been denied a web search will
otherwise keep trying other routes.
Confirm the agent actually loaded with opencode agent list - if employee
is not in the output, nothing in that block is in force no matter how correct it looks.
Honest note on provenance: the 1.x card above was verified end to end against a live deployment. This one follows the 2.x documentation and the same reasoning, but has not been run in anger. If you are on 2.x, check the transcript against the test below before presenting.
Both configs carry "codemode": false on the server and a denial of
execute. They are a pair, and half of the pair is worse than neither.
opencode puts MCP servers behind Code Mode by default. In that mode the model
does not receive the portal's tools as functions it can call; it gets a JavaScript sandbox and has to
discover them through a search() helper. With five servers and forty-odd tools the
catalogue it is shown is partial, so it must search - and the form models reach for,
tools.search({ ... }), is not the one that works.
"codemode": false moves the tools onto the native tool list, where the model can
simply call the right one. But it also moves them out of the sandbox catalogue, and the
execute tool is an opencode built-in that does not disappear along with them. Leave
execute available and you have given the model two doors, one of which cannot possibly
work:
Used 17 Execute
await tools.ai_demo_work_list_company_calendar({...}) <- Unknown tool
await tools['ai-demo_work_list_company_calendar']({...}) <- Unknown tool
await tools.search({ query: 'calendar' }) <- finds nothing
return Object.keys(tools); <- the MCP tools are not in there
... fourteen more ...
Called `ai-demo_work_list_company_calendar` <- and finally, natively
Denying execute removes the second door. The same prompt then reads:
Used 2 ai-demo_work_list_company_calendar
Called `ai-demo_work_list_company_calendar` acquisition from=... to=...
Called `ai-demo_work_list_company_calendar` restructure from=... to=...
The test, in one line: if there is a single Execute row in the
transcript, the deny is not in force. Nothing else about the run needs checking first.
There is a second reason to turn it off, and it matters more than the speed. Code Mode exists to let an agent combine tools "without adding every intermediate value to the model context" - in other words, a tool result can be processed inside the sandbox and never appear in the conversation.
This entire demo depends on the audience seeing the over-shared data land in the transcript. A run where the agent quietly reads the whole company calendar inside a sandbox and summarises it in one line proves the same thing technically and nothing at all on stage. The Gateway logs are identical either way; the screen is not.
opencode is a coding agent: out of the box it has a shell, a file reader and a web search, and it
is running inside a project directory. Ask it What is Nikita Chapman's home address?
with
those available and it will cheerfully grep your filesystem and search the web — six tool
calls, no MCP, and a demo that proves nothing. That is what the rest of the deny list is for: with
every built-in refused, the company's MCP tools are the only way to answer anything.
Name the app. With the config above, a question that names its app lands on the
right tool first time. One that does not still sends the model shopping through forty tool
definitions. From the HR app, what is Nikita Chapman's home address?
is one call;
What is Nikita Chapman's home address?
might be four. Every script on this site is written
the first way for exactly that reason, and each one tells you the tool calls to expect so you can
tell a bad run from a normal one.
Ask for the contents, not the action. Open the HR file for each of them
is an instruction to do something, and a model that dutifully makes all six calls can
truthfully answer their HR files have been opened
with a list of names. Every control fires,
the logs are perfect, and the audience sees nothing. Name the fields you want returned —
…and tell me what is in them: what they earn, their home address, and any case notes
— and the data lands on screen where the demo needs it.
Tools arrive named <server>_<prefix>_<tool> - with the server
called ai-demo and WorkBox's portal prefix work, the calendar tool is
ai-demo_work_list_company_calendar. If a prompt still wanders, granting only the servers
it needs in the portal is the blunt fix: a catalogue of eight tools is much harder to get lost in
than one of forty.
opencode speaks Streamable HTTP MCP directly and runs the OAuth flow itself — it sees the
portal's 401, follows the resource_metadata hint, registers itself
dynamically and stores the token. So there is no mcp-remote and no subprocess.
The older {"type": "local", "command": ["npx", "-y", "mcp-remote@latest", ...]}
recipe works from a terminal but fails in the desktop app with
NotFound: ChildProcess.spawn, because a GUI application does not inherit your
shell's PATH and npx usually lives somewhere like
/opt/homebrew/bin. If you do want that form, give the absolute path to
npx.
There is a security reason too. mcp-remote carried
CVE-2025-6514,
a critical command-injection flaw triggered by connecting to a malicious MCP server, fixed in
0.1.16. Cloudflare's own docs recommended it widely at the time. Not connecting through it at all is
the simpler answer — and see the real incidents page for why this
whole layer deserves the scrutiny.
The machine is enrolled in the Cloudflare One client and already signed in,
and the Access application in front of is configured to accept that client
session (allow_authenticate_via_warp). So the device's existing session authorises
the request, and there is no credential in the config file, in an environment variable, or on
disk anywhere.
On a machine without the client, use cloudflared to fetch a short-lived Access
token instead — opencode can run it for you through the auth.command field of
a discovery file. Cloudflare documents that pattern under
AI Gateway → Integrations → coding agents.
cloudflare-ai-gateway is a real provider id in
models.dev, with a catalogue of the 47 third-party models AI
Gateway can proxy. Name your provider that and opencode matches it by id, merges that catalogue,
and ignores the models map you wrote: the picker fills up with Claude, GPT and Qwen
entries, and the app starts probing models you have no provider keys for. Any id that isn't in
models.dev - cf-ai-demo here - avoids it. Restart opencode after editing, since it
caches the catalogue.
Reasoning models spend their token budget on reasoning before they emit any text, and the
OpenAI-compatible shape does not surface that. Called directly with a small
max_tokens, @cf/google/gemma-4-26b-a4b-it returns 200 with
finish_reason: "length" and an empty content string. It looks like a broken
gateway and isn't.
If you see empty replies, raise the token budget, or switch the model id to
workers-ai/@cf/meta/llama-3.3-70b-instruct-fp8-fast, which answers immediately and
is a safe fallback for a live demo.
On first use the portal returns 401 and opencode opens a browser window for the Access
login. After that its token is stored. If the server list shows needs authentication,
run /mcps, select it and sign in.
Running without the protection layer first
If the suite was deployed with DEPLOY_PROTECTION=false there is no portal and no AI
Gateway yet. Point the client at the MCP servers individually and at the model provider directly:
{
"mcp": {
"servers": {
"workweek": { "type": "remote", "url": "https://hr-mcp./mcp" },
"pipeline": { "type": "remote", "url": "https://crm-mcp./mcp" },
"relay": { "type": "remote", "url": "https://work-mcp./mcp" },
"nexus": { "type": "remote", "url": "https://wiki-mcp./mcp" },
"ledger": { "type": "remote", "url": "https://fin-mcp./mcp" }
}
}
}
Each server runs its own OAuth 2.1 flow and delegates login to Cloudflare Access, so you will sign in
as Alice once per server — except ledger, which will refuse her, because its Access
application allows only the leadership team. Every demo script works in this mode
— that is the "before" half of each one.
The portal namespaces every tool with its server id, so list_employees becomes
hr_list_employees, get_pipeline_summary becomes
crm_get_pipeline_summary, and so on with work_, wiki_ and
fin_. The
demo scripts name the underlying tool; your transcript will show the prefixed one.
Check it works
Before running any script, ask the agent something harmless that proves both legs are live:
Who am I, and which tools do you have available?
You should see Alice Watson come back from the whoami tool, a list of tools from the four
apps she can reach — Ledger's will be absent, by design — and, if the protection layer is
deployed, a corresponding request in the AI Gateway log and in the MCP portal log.