Agent MCP server
1. Foundational mental model
Section titled “1. Foundational mental model”The Bun game server already holds every live room in memory, keeps country stats in SQLite and writes a seven-day diagnostics database. The MCP endpoint is a read window onto those three stores for an agent such as Cursor. An agent sends JSON-RPC over HTTP POST to /mcp with Authorization: Bearer <token>, and the server answers with one JSON text block per tool call.
Each request is independent. The handler builds a fresh transport and a fresh McpServer for every POST, so there is no MCP session, no server-sent event stream and nothing to clean up when an agent disconnects.
The tools split by the store they read. Only four sources exist, and none of them is written by a tool call.
2. Implementation status & known gaps
Section titled “2. Implementation status & known gaps”The endpoint ships in the game server and is used against local and production hosts through .cursor/mcp.json. It is off by default. Setting MCP_ENABLED=1 (or true) without MCP_AUTH_TOKEN stops the process at boot, and a token shorter than 32 characters in production (16 elsewhere) does the same. A successful boot logs MCP enabled at /mcp (Streamable HTTP, bearer auth).
The server registers tools only. There are no registerResource or registerPrompt calls, so resources/list and prompts/list return nothing useful to an agent.
3. Concrete implementation
Section titled “3. Concrete implementation”Routing
Section titled “Routing”/mcp is a static route on the same Bun.serve instance as the WebSocket gateway described in Bun WebSocket gateway. All four methods are declared so the handler can answer GET and DELETE with a JSON-RPC 405 instead of Bun’s default.
const server = serve({ hostname: env.SERVER_LISTEN_HOST, port: env.PORT, routes: { "/mcp": { OPTIONS: () => handleMcpOptions(), GET: (request) => handleMcpRequest(request), POST: (request) => handleMcpRequest(request), DELETE: (request) => handleMcpRequest(request), },Request handler
Section titled “Request handler”The order of checks matters. Disabled deployments answer a plain 404 so the endpoint is not advertised. The rate limit is consumed next, before the bearer check, so each failed guess costs budget.
if (access === "disabled") { return new Response("Not Found", { status: 404 }); }
const clientIp = getClientIPAddress(request) ?? "unknown"; const rateLimit = rateLimiter.consume("MCP_REQUEST", { ipAddress: clientIp }); if (!rateLimit.allowed) { return mcpJsonRpcError(429, "Too many requests"); }
if (access === "unauthorized") { return mcpJsonRpcError(401, "Unauthorized"); }
// Stateless mode: no session-scoped GET/DELETE streams. Clients treat 405 as optional SSE. if (request.method === "GET" || request.method === "DELETE") { return mcpMethodNotAllowed(); }
const transport = new WebStandardStreamableHTTPServerTransport({ sessionIdGenerator: undefined, }); const server = createFlagsMcpServer();
try { await server.connect(transport); return await transport.handleRequest(request); } catch (error) { logger.error("MCP request failed", error); return mcpJsonRpcError(500, "Internal server error"); }sessionIdGenerator: undefined puts the SDK’s WebStandardStreamableHTTPServerTransport in stateless mode: it issues no mcp-session-id and answers each POST in its own response. The CORS preflight returns 204 with Access-Control-Allow-Methods: GET, POST, DELETE, OPTIONS and allows the Authorization, Content-Type, mcp-session-id, mcp-protocol-version and Last-Event-ID headers. No Access-Control-Allow-Origin header is set.
Bearer check
Section titled “Bearer check” if (!authHeader?.startsWith("Bearer ")) { return false; }
const providedToken = authHeader.slice("Bearer ".length).trim(); if (providedToken.length === 0) { return false; } if (providedToken.length !== expectedToken.length) { return false; }
try { return timingSafeEqual(Buffer.from(providedToken), Buffer.from(expectedToken)); } catch { return false; }The scheme is case-sensitive: bearer abc fails. The same secret also authorizes POST /api/internal/diagnostics, which the SvelteKit app uses to forward daily completions into the diagnostics database.
Boot validation
Section titled “Boot validation”if (env.MCP_ENABLED) { if (!env.MCP_AUTH_TOKEN) { throw new Error("MCP_AUTH_TOKEN must be set when MCP_ENABLED=1"); } validateMcpAuthTokenLength(env.MCP_AUTH_TOKEN, env.NODE_ENV);}Tool registration
Section titled “Tool registration”Every tool returns its payload through one helper, pretty-printed into a single text content block. Agents parse the JSON themselves; there is no structuredContent and no outputSchema.
const jsonToolResult = <T>(data: T) => ({ content: [{ type: "text" as const, text: JSON.stringify(data, null, 2) }],});
const parseSince = (value: string | undefined): number => { if (!value) { return Date.now() - 7 * 24 * 60 * 60 * 1000; } const parsed = Date.parse(value); return Number.isFinite(parsed) ? Math.max(parsed, Date.now() - 7 * 24 * 60 * 60 * 1000) : Date.now() - 7 * 24 * 60 * 60 * 1000;};parseSince clamps every diagnostics query to the retention window, so an agent asking for last month gets the last seven days. An unparsable timestamp falls back to the same window instead of failing.
The fifteen tools
Section titled “The fifteen tools”| Tool | Inputs | Source | Purpose |
|---|---|---|---|
server_health | none | process | Liveness check with a UTC timestamp. |
server_stats | none | memory | Counts of rooms, connected users and active games. |
list_rooms | limit 1 to 200, default 50 | memory | Summaries of live rooms; count is the total, returned the page. |
get_room | inviteCode | memory | One room summary by invite code. |
get_room_game_state | inviteCode | memory | Timers, leaderboard, members, settings and deletion flag for one room. |
replay_solo_ledger | sessionToken, ledger | ledger manager | Validates and scores a signed solo ledger without persisting; returns persisted: false. |
generate_daily_session | date (YYYY-MM-DD) | shared generator | The deterministic daily rounds for a UTC date, with answers. |
generate_solo_questions | difficulty, regionFilter, count 1 to 250 | shared generator | A sample solo run from the production generator, random order per call. |
leaderboard_countries | difficulty, gameType, limit 1 to 100, default 25 | stats.db | Country accuracy leaderboard for one stratum. |
recent_matches | since, limit 1 to 100, cursor | diagnostics | Paged, sanitized match summaries. |
get_match_diagnostics | diagnosticId (UUID) | diagnostics | One match summary plus its events. |
get_match_timeline | diagnosticId, limit 1 to 100 | diagnostics | Ordered lifecycle and exception events for one match. |
incident_timeline | since, until, category, severity, code, limit 1 to 200, cursor | diagnostics | Filtered operational and security events. |
security_summary | since, until | diagnostics | Aggregated security and rate-limit counts. |
forwarding_status | since, until, diagnosticId, limit 1 to 200 | diagnostics | Outcomes of stats, Convex, replay and XP forwarding. |
Missing rooms and invalid dates come back as a normal result such as { "error": "Room not found", "inviteCode": "…" }, with isError unset. Input validation failures from the Zod schemas are reported by the SDK.
Answer masking in room snapshots
Section titled “Answer masking in room snapshots”function snapshotQuestion(question: GameQuestion, phase: GamePhase) { const base = { index: question.index, countryCode: question.country.code, optionCodes: question.options.map((country) => country.code), startTime: question.startTime, endTime: question.endTime, };
if (revealCorrectAnswers(phase)) { return { ...base, correctAnswer: question.correctAnswer }; }
return base;}The mask removes one field. countryCode in base names the flag being asked about, so an agent can read the answer during the question phase. This is the same leak the player wire has, covered in Room lifecycle and host migration.
One tool call end to end
Section titled “One tool call end to end”4. Internal mechanics & mathematics
Section titled “4. Internal mechanics & mathematics”Rate limit
Section titled “Rate limit”MCP_REQUEST uses the same sliding-window counter as the WebSocket gateway, with rule (L = 60) per (T = 60,000) ms keyed by IP. At time (t) into the current window, with (c) requests counted in it and (p) in the previous window:
A denied request does not increment (c). The long-run ceiling is therefore about one request per second per IP, which is ample for an agent that issues a handful of tool calls per task.
Brute-force cost
Section titled “Brute-force cost”openssl rand -base64 32, the generator the reference doc suggests, produces 44 characters carrying 256 bits. Because the length check returns early, an attacker can learn the length, which removes no entropy from a random token. At 60 guesses per minute from one IP, the expected time to hit the token is
The production floor of 32 characters leaves at least (32 \times 6 = 192) bits when the characters are random base64, which is still out of reach. The floor checks length only, so a 32-character dictionary phrase passes validation with far less entropy.
Per-request cost
Section titled “Per-request cost”Each POST constructs an McpServer and registers 15 tools with their Zod schemas before handling a single JSON-RPC message. The work is proportional to the tool count and independent of room count, and the rate limit caps it at roughly one construction per second per IP. The tools themselves are bounded by their limit arguments: at most 200 room snapshots, 250 generated questions, or 200 diagnostics rows per call.
5. Threat model, failure modes & edge cases
Section titled “5. Threat model, failure modes & edge cases”| Case | What happens |
|---|---|
| MCP disabled | GET, POST and DELETE return 404. OPTIONS still returns 204 with the CORS headers, so a preflight reveals that the route exists. |
| Missing or wrong token | 401 after a rate-limit slot is spent. |
| Leaked token | Full read access, including live answers and future daily rounds. Rotation is manual: change MCP_AUTH_TOKEN and restart. The same rotation breaks daily diagnostics forwarding until the web app is updated. |
| Spoofed proxy headers | getClientIPAddress trusts cf-connecting-ip, true-client-ip and x-real-ip. If the API host is not behind Cloudflare, a caller can rotate these headers to get a fresh 60-per-minute bucket on each request. |
| No trusted header | All callers share the unknown key, so any traffic to /mcp can exhaust the operator’s budget. |
| Agent opens an SSE stream | GET returns 405; the SDK client treats this as “no stream” and continues with POST. |
| Tool throws | The handler logs the error and returns 500 with JSON-RPC code -32000. |
| Room deleted between calls | The next call returns { "error": "Room not found" }. Snapshots are point-in-time. |
| Ledger oracle | replay_solo_ledger reports whether a ledger’s signature and scoring pass without recording a rejection. A token holder can iterate on forged ledgers without tripping the rejection counters that the public route updates. |
6. Architectural tradeoffs & non-goals
Section titled “6. Architectural tradeoffs & non-goals”Use this endpoint when an agent needs to inspect a live room, reproduce a generator output or read recent incidents without shell access to the host.
Use something else when the task changes state. Admin actions belong behind ADMIN_API_KEY routes with human review, as the reference doc states.
The reference doc gives the reasons for skipping Arcjet and for keeping the tools read-only. The other rationales below are this chapter’s reading of the code.
| Choice | Alternative | Why this one |
|---|---|---|
| Stateless transport, new server per request | Session IDs with a server-side map | No session store to expire or leak; the cost is rebuilding 15 tool definitions per call. |
| Shared bearer secret | OAuth or per-user tokens | One operator and a few agents; OAuth would need an authorization server for no extra audience. |
| Rate limit before auth | Auth first, limit only valid tokens | Guessing costs budget; the price is that anonymous traffic can crowd out the operator on a shared IP. |
No Arcjet on /mcp | Route through withMiddleware | Avoids bot-detection false positives on agent user agents (stated in the reference doc). |
| JSON text content | structuredContent with output schemas | Works with every MCP client; agents lose typed results. |
| In-process with the game server | Separate diagnostics service | Direct access to in-memory rooms, which no other process can see. |
Non-goals: mutating rooms, kicking players, issuing XP, exposing MCP resources or prompts, streaming notifications, and multi-tenant access control.