Adding a Service¶
The platform is designed so that adding the Nth service requires no code changes — only configuration. The five steps below are the complete procedure.
Adding a new Identity Provider¶
If your new service needs its own credential-linking flow (rather than
reusing one already configured), add an entry to broker.identityProviders
in your HelmRelease values. Each entry's alias doubles as the id shown on
the portal's Identities page — no separate mapping to keep in sync.
type: oauth21-direct— use this when the service is itself an OAuth 2.1 authorization server (e.g. rucio-mcp). No Keycloak IdP configuration is needed at all; the broker is a direct OAuth 2.1 client via its own CIMD document (GET /.well-known/cimd). Also requiresbroker.publicOriginto be set to the portal's origin (see below) — the broker refuses to start otherwise. See Rucio: Per-Site Setup for a concrete, deployed worked example of this provider type (one entry per Rucio site).
broker:
publicOrigin: "https://mcp-portal.af.uchicago.edu"
identityProviders:
- type: oauth21-direct
alias: my-service-oauth
targets: ["my-new-service"]
authorizationEndpoint: "https://my-new-service.example/authorize"
tokenEndpoint: "https://my-new-service.example/token"
issuer: "https://my-new-service.example"
displayName: "My New Service"
enables: "Access to my-new-service on your behalf"
broker.publicOrigin is the canonical origin (scheme + host, no trailing
slash) every OAuth 2.1 URL the broker constructs itself is built from: the
redirect_uri it sends to the service's authorization server, and every
redirect_uris entry in the CIMD document above. Register the service's
authorization server client with exactly
<publicOrigin>/v1/oauth/callback/<alias> as its redirect_uri whitelist
entry (or point it at the CIMD document, which advertises the same URL).
type: keycloak-brokered— use this when the service is (or can be registered as) an OIDC identity provider. This does require configuring the service as an Identity Provider in Keycloak (Settings → Identity Providers), with "Store Tokens" and "Stored Tokens Readable" both on, plus theread-tokenclient role from Keycloak'sbrokerclient granted to callers (seedocs/auth.md).
broker:
identityProviders:
- type: keycloak-brokered
alias: my-service-oidc
targets: ["my-new-service"]
displayName: "My New Service"
enables: "Access to my-new-service on your behalf"
type: x509— use this when the service's real credential is a VOMS proxy (e.g. ami-mcp). Unlike the two types above, delivery is service-side redemption, not header injection: the aggregator injects only an AF Broker Identity Token, and the service redeems the caller's proxy itself viaPOST /v1/credentials/x509/redeem(issue #112's wire format — proxy PEM material never transits the aggregator). The service must also be markedauth_type: x509in the aggregator service list (Step 1); the broker refuses to start when the entry'stargetsand theauth_type: x509services drift in either direction — including when a service has no explicit entry at all, since there is no synthesized fallback. WithserviceUrlset it also requires the broker signing key (broker.identityToken.existingSigningKeySecret) and the shared Vault connection (broker.oauth21.tokenStore.vault) — omitserviceUrlfor the legacy ephemeral-Job mint path (signing key omitted there just warns). This entry replaces the removed globalbroker.env.VOMS_TOKEN_SERVICE_URL— see x509 deployment notes.
broker:
identityProviders:
- type: x509
alias: x509
targets: ["my-x509-service"]
serviceUrl: "http://voms-token-service.voms-token.svc.cluster.local:8080"
voms: "atlas" # optional, default "atlas"
valid: "192:00" # optional, default "192:00"
displayName: "Grid certificate (x509)"
enables: "VOMS proxy minting for x509-authenticated services"
See docs/auth.md#identity-provider-types for how the types differ.
Migrating from the pre-unification chart values¶
Older chart releases configured identity providers across four separate
values. All four are consolidated into broker.identityProviders above:
| Old value | New equivalent |
|---|---|
broker.oidc.idpAlias |
A keycloak-brokered entry's alias |
broker.oauth21.providers |
oauth21-direct entries (same fields, still camelCase) |
broker.cimd.idpAliases |
Derived automatically from oauth21-direct entries — remove this value entirely |
broker.identitiesLinkClientId |
Removed entirely — keycloak-brokered entries no longer need it; the portal links them via its own client-side flow regardless |
Step 1 — Add the service to the aggregator service list¶
Edit the HelmRelease for the platform (typically
clusters/<cluster>/af-mcp-platform/helmrelease.yaml) and add one entry under
values.aggregator.services:
values:
aggregator:
services:
# existing services omitted for brevity
- name: my-new-service
url: http://my-new-service.af-mcp-backends.svc.cluster.local:8000/mcp
required_permission: my-new-service:use
timeout_seconds: 30
auth_type: none # or "bearer" (default) / "x509" -- see below
tools_cache_ttl: 300 # seconds; see below
required_permission is the permission string the broker's Authorization
subsystem will check before forwarding any tool call to this service.
When adding a service, also state its trust tier (Elwood v5 / Shannon:
user-tier, service-tier, or infrastructure-tier) in its deployment
manifests in the GitOps repo, and choose required_permission accordingly
— see docs/architecture.md "Trust tiers". required_permission: __none__
(open to any authenticated user) is only appropriate for user-tier
read-only services.
auth_type controls what per-user credential the aggregator injects into
the service call and defaults to bearer if omitted — meaning the
broker will try to mint a per-user credential for every caller, which
requires an identity provider configured for this service's name (see
"Adding a new Identity Provider" above) and fails with a friendly
"not linked" error otherwise. Set auth_type: none explicitly if the
service authorizes itself some other way (e.g. a platform k8s service
account) and needs no per-user credential forwarded at all. auth_type: x509
marks a service whose per-user credential is a VOMS proxy (e.g. ami-mcp):
the aggregator injects an AF Broker Identity Token (aud = the service's
audience, defaulting to its name — issue #257) and the service
redeems the caller's cached proxy itself via
POST /v1/credentials/x509/redeem (issue #112) — this requires the broker
signing key to be mounted (broker.identityToken.existingSigningKeySecret),
and the service to run in a mode that verifies broker JWTs and redeems
proxies (ami-mcp's --auth broker, via the af-credentials library). The new service's name must be added to some identityProviders entry's
targets — every auth_type: x509 service needs an explicit entry
covering it, or the broker refuses to start naming it (there is no
synthesized fallback; see "Adding a new Identity Provider" above).
For a bearer service, the aggregator also attempts a best-effort per-user
credential mint during tools/list (not only tools/call), so a service
whose own MCP endpoint requires auth just to list tools (e.g. rucio-mcp)
isn't invisible to every caller — see mcp/aggregator.py's
_make_client_factory docstring for the full rationale.
tools_cache_ttl (seconds, default 300, matching fastmcp's own
ProxyProvider default) controls how long a cached component list for this
service may be served from a by-name lookup (e.g. resolving a tool during a
tools/call) before refreshing. Tool schemas are assumed
caller-independent, so a schema cached under one caller's credential being
served to another isn't a credential leak — only set this to 0 for a
service whose tool list genuinely personalizes per caller.
apply_namespace — tool naming¶
apply_namespace controls whether the aggregator mounts this service's
tools as <prefix>_<toolname> and defaults to true — the safe
choice, since it's what prevents two services from advertising the same
tool name and one silently shadowing the other in tools/list. Leave it
unset unless you have a specific reason to change it.
Set it to false only for a service whose tools are already self-prefixed
at the source (baked into the tool names the service itself advertises,
not added by the aggregator). rucio-mcp is the shipped example: it serves
tools already named rucio_list_dids, rucio_whoami, etc., so leaving
apply_namespace: true would double-prefix them into
rucio_rucio_list_dids. The shipped services.yaml therefore sets
apply_namespace: false on its rucio entry, and callers see the plain
rucio_list_dids name. The aggregator routes calls to an apply_namespace:
false service by that same prefix, so every tool it advertises must be
named <prefix>_<tool> -- a tool name missing the prefix is unreachable
even though it can still appear in tools/list.
false is only safe when no other configured service can advertise an
overlapping tool name — with one rucio site configured, that holds. It
stops holding the moment a second self-prefixed service enters the picture:
configuring both an ATLAS and an ESCAPE rucio site as separate services
(see the shipped services.yaml comment) means both would advertise the
same un-namespaced rucio_* names, and fastmcp resolves un-namespaced
mounts in registration order — the second one silently shadows the first
instead of failing loudly. The accepted fix for that case is to set
apply_namespace: true on both site entries and accept the resulting
double-prefixed names (rucio_atlas_rucio_whoami, rucio_escape_rucio_whoami)
— ugly, but unambiguous and requires no upstream rucio-mcp change. See
#113 for the
full tradeoff discussion.
Whatever you choose, the deployed tool names are what callers actually see
via tools/list on /mcp — confirm the names you expect show up there (see
Verification below) rather than assuming from the config alone. GET
/v1/catalog and the portal's Catalog page show the service itself (name,
permission, auth type); the individual tool names live one level down at
GET /v1/catalog/{service}/tools, which the portal fetches when a server
card's Tools section is expanded.
exclude_tools — hiding backend tools entirely¶
exclude_tools is a list of native (pre-namespace) tool names a backend
should never be exposed for, at all — not "gated behind a permission", but
absent from tools/list and refused on tools/call exactly as if the
backend never advertised them:
- name: condor_service
prefix: condor
url: "http://condor-mcp:8000/mcp"
required_permission: manage_jobs
exclude_tools: ["advertise_to_collector"]
Use this for a backend tool that's infrastructure-facing with no user story
behind the broker (condor-mcp's advertise_to_collector, the shipped
example), not for something you merely want fewer callers to reach — that's
what per-tool required_permission (Step 2 below) is for. Names are
whatever the provider itself sees, which follows apply_namespace the same
way required_permission's dict keys do: the backend's own name for the
(default) namespaced case, or the raw name including its self-declared
prefix when apply_namespace: false. A name that never appears in the
backend's own listing logs a one-time aggregator.exclude_tools_not_found
warning — almost always a typo.
Naming conventions¶
New services are named <backend>_service per the Elwood v5 glossary (e.g.
rucio_service). Methods (MCP "tools") use verb_noun naming (e.g.
list_dids, submit_job). This documents the convention for new services —
it renames nothing that already exists.
name is the registry identity, not the wire audience. For AF-native
(x509/broker-issued) services, the broker mints an AF Broker Identity Token
whose aud each backend validates. That aud is the service's audience
field, defaulting to name — so set an explicit audience and you can adopt
the <backend>_service name (registry key, catalog, audit label) while the
backend keeps validating its historical audience. Renaming name without an
explicit audience silently re-points the contract and 401s the backend
(learned the hard way, 2026-08-26 — see docs/auth.md's "AF Broker Identity
Token"). A rename is: name: <backend>_service + audience: <old-name>, with
the identity-provider targets updated to the new name in lockstep.
Reserved: the af prefix and the builtin gateway service name. The
registry always carries a builtin gateway service (issue #240) — the
gateway's own identity, catalog, and usage methods (af_whoami,
af_list_identities, af_list_mcp_servers, af_link_identity, af_usage),
served by the aggregator itself rather than proxied to any backend. Its name
defaults to gateway_service and is deployment-configurable
(broker.builtinServiceName / BUILTIN_SERVICE_NAME); the reserved af
prefix is fixed, since the method wire-names are built from it. It is not a
services.yaml entry and cannot be one: an entry claiming the af prefix or
the configured builtin name fails registration with a clear error, since
either would let a configured service shadow (or replace) the methods a caller
relies on precisely when everything else is broken.
Step 2 — Pick (or reuse) a permission for the service¶
required_permission in services.yaml (Step 1 above) is the sole
declaration of what a service target requires — the service registry, not
policy.yaml, is authoritative here (see issue #60). policy.yaml's only
remaining job is mapping permissions to Keycloak groups via
group_permissions (Step 3 below).
required_permission has four forms:
- A permission name (e.g.
read_data) — the caller must hold that permission, granted viagroup_permissions, for every tool this service exposes. __none__— open to any authenticated user; no permission needed. Use this only as a deliberate, explicit opt-in.- Omitted entirely — no permission gate; the credential layer becomes
the gate instead (the caller must have a linked identity / mintable
credential for this target, which is itself the authorization). The
broker refuses to start if a service omits
required_permissionand no credential provider resolves for its target either (e.g.auth_type: bearerwith noidentity_providersentry naming it, orauth_type: nonewith nothing registered) — that combination would mean the service has no gate at all, neither a permission nor a credential requirement. - A dict, keyed by tool name — different tools of the same service can require different permissions:
required_permission:
__default__: manage_jobs # optional -- see below
query_jobs: read_monitoring
query_history_db: read_monitoring
Keys are the tool's native name — the name the backend itself
advertises, before the aggregator's <prefix>_ namespacing is applied (or
the raw name, unchanged, if apply_namespace: false) — so the same key
works regardless of that choice. __default__ is the permission any tool
not listed as its own key falls back to, and it's optional, not
required: if you omit it, an unlisted tool is disabled outright (nobody
can call it, regardless of what they hold) rather than silently inheriting
some other tool's permission or falling open. This is deliberate —
per-tool gating is opt-in, so a new tool a backend adds later can never
slip through under a permission nobody meant to grant it. Give every tool
you want reachable its own key, or set __default__ to whatever you want
the safe fallback to be — usually the service's most privileged
permission, with only the genuinely read-only tools de-escalated by their
own key, exactly as the example above does for query_jobs/
query_history_db. A service can appear as a target under more than one
permission's "your access" grants (GET /v1/permissions) when it uses
this form, and
GET /v1/catalog/{service}/tools reports each visible tool's own resolved
permission — GET /v1/catalog's single permission field is null for a
dict form with no __default__, since there's no one value to summarize.
If a backend renames or adds tools, the dict and the backend's listing
drift apart: unmapped tools (advertised, no key, no __default__) are
silently disabled, and stale keys map nothing. The broker logs
entitlement.tool_mapping_drift once per change, counts them in
af_mcp_tool_mapping_drift_total, and lists them at
GET /v1/admin/tool-mapping-drift (the portal's admin page), as observed
when a caller lists the service's tools.
Permission names must be built-in or declared. Besides the built-in
permissions (read_data, submit_jobs, ... see PERMISSIONS in
authorization/base.py), a site can invent its own, but it must declare the
permission's action type (read or state_change) so the catalog and audit
labels are right:
The broker refuses to start if a service requires a permission that is
neither built-in nor listed in custom_permissions (an unknown name could
otherwise only be labelled read). Declaring it does not grant it to anyone;
Step 3 still applies.
Overriding a tool's action type. A tool's action_type (read or
state_change) comes from its permission. To override it for specific tools,
use entitlements.target_action_types. Keys are service names (the
name in services.yaml, not the prefix), and the globs match the wire
tool name a caller sees, including the <prefix>_ namespace unless the
service sets apply_namespace: false:
A key naming no registered service, or a glob matching no tool, silently overrides nothing.
If an existing permission already covers the new service (e.g. a generic
read_metadata that several services already require), reuse it and skip to
Step 4 — no policy change needed. Only continue to Step 3 if you're
introducing a genuinely new permission name.
Step 3 — Map the permission to Keycloak groups (if new)¶
If you added a new permission in Step 2, map it to one or more AF Keycloak groups in the HelmRelease values:
values:
entitlements:
group_permissions:
# existing mappings omitted
af-my-new-service-users:
- my-new-service:use
Principals in the af-my-new-service-users Keycloak group will be granted
my-new-service:use. If you forget this step (or typo the permission name),
the broker refuses to start, naming both the service and the permission
it can't reach — see docs/auth.md's "Group-to-Permission Mapping Example".
Step 4 — Allow egress to the service in NetworkPolicy (if needed)¶
The broker's NetworkPolicy allows in-cluster egress to service pods in the same namespace as the broker, on a configured list of ports. If your new service is in the same namespace and uses one of the default ports (8000, 8080), no change is needed.
If it listens on a different port, append it to
networkPolicy.broker.servicePorts in your HelmRelease values — no template
edit needed:
Cross-namespace services are not currently exposed through values — the
namespace selector in templates/networkpolicy.yaml is scoped to the release
namespace. To reach a service in a different namespace you'd need to either
move it into the broker's namespace or extend the chart's egress rule. If
that becomes a recurring need, file an issue to parameterize
serviceNamespaces similarly.
Step 5 — Redeploy¶
Flux will render the new HelmRelease values, update the aggregator ConfigMap, and
roll the broker pods. The new service appears as its own entry in
GET /v1/catalog once the pods are healthy; its individual tool names show up
in tools/list over /mcp (see Verification below).
Tool-authoring conventions¶
Steps 1–5 wire a service in with no code. This section is about the code you
do write — the tools the backend itself advertises. The aggregator forwards
whatever the backend serves verbatim, so the quality of each tool's
description, error text, and schema is set entirely at the source. These five
conventions cost every LLM agent a wasted retry cycle when skipped, regardless
of which client calls the tool or how good its own discovery is, so treat them
as the bar for every backend — existing and new. rucio-mcp already meets the
first two (one-line summaries and actionable errors) and is cited below as
the shipped example for those; outputSchema and annotations are
documented here but, as of this writing, implemented by no backend in the
fleet yet — maniaclab/af-mcp-platform#238 is the rollout tracking issue.
The broker's own gateway_service tools
(broker/src/af_mcp_broker/mcp/diagnostics.py) declare both already, and
are the in-repo reference for outputSchema and annotations, though not
for the Annotated[CallToolResult, Model] markdown-preservation pattern
below, which none of their tools need (they return plain pydantic models,
no curated markdown to keep alongside the structured payload).
-
Lead each tool's description with a one-line summary. The first line is what a client shows in a compact tool list and what it ranks on. This matters more as clients move from loading every tool's full schema upfront to progressive tool-search over just the summaries, where a tool whose first line doesn't say what it does never gets picked. Put the one-liner first, then the detail (arguments, caveats, when to call it) below it. The builtin
gateway_servicetools follow this — seeaf_whoami's "Return the caller's own subject, groups, and effective permissions." ahead of its usage notes inbroker/src/af_mcp_broker/mcp/diagnostics.py. -
On an empty or error result, return an actionable next step, not just the raw condition. "No replicas found" or a bare backend exception makes the agent guess (and usually burn a diagnostic call, or several — the frictions in #216 came from exactly this). Say what to do instead. rucio-mcp is the model: several of its tools' empty/error results carry a "Next steps" hint, e.g. on an empty dataset-replicas result it points the caller at "if this is a container DID, use
rucio_list_container_replicasinstead". Name the specific tool or argument that fixes the situation whenever you can. Surface it as a structured MCP error (isError: truewith the hint in the message), not raw backend exception text. -
Declare
outputSchemawherever the backend's SDK supports it, alongside the curated markdown, not instead of it. Without an output schema a client can't type-check a tool's result or programmatically compose one tool's output into another's input; it's left parsing prose. Note that an MCPoutputSchemaMUST be a JSON object schema — a tool that returns a bare list gets no usable schema (or a synthetic single-key wrapper with a meaningless field name), so wrap a list in a small object model with a named field. Annotating a tool-> MyModeldirectly gets yououtputSchema -
structuredContent, but the SDK then renders the text block as the raw JSON dump of that model, destroying any curated markdown you wrote. The supported escape hatch isAnnotated[CallToolResult, ResultModel]: return an ordinaryCallToolResultwith your ownTextContentandstructured_content=payload.model_dump(mode="json"), and the SDK still publishesResultModel's schema and validatesstructured_contentagainst it — the annotation is metadata for schema/validation purposes only, never evaluated as the literal return type.Context,CallToolResult,TextContent,ToolAnnotations, and every model used in a return annotation must stay runtime imports, neverTYPE_CHECKING— the SDK'sfunc_metadata()resolves live annotations viainspect.signature(func, eval_str=True), and aTYPE_CHECKING-only name raisesInvalidSignatureat registration. Thegateway_servicetools are the in-repo reference instance for the schema itself: all five declare an objectoutputSchema, andaf_list_identities/af_list_mcp_serversshow the list-wrapping pattern (ListIdentitiesResult.identities/ListMcpServersResult.servers) inbroker/src/af_mcp_broker/mcp/diagnostics.py— none of them need theAnnotated[CallToolResult, Model]form since they have no curated markdown to preserve alongside the structured payload. -
Declare read-only/destructive annotations on every tool.
ToolAnnotations(read_only_hint,destructive_hint,idempotent_hint,open_world_hint,title) tells a client which tools are safe to call speculatively and which change real state, without it having to guess from the name or description. Setread_only_hint=Truefor anything that only reads;destructive_hint/idempotent_hintare meaningful only whenread_only_hintisFalse, so leave them unset on a read-only tool. The broker forwards annotations verbatim and lints them againstpolicy.yaml's resolved action type (maniaclab/af-mcp-platform#238B.8) — a disagreement doesn't block anything (the gateway never trusts a backend's own annotation for an authorization decision), but it does show up on the admin page as something worth reconciling. Thegateway_servicetools (broker/src/af_mcp_broker/mcp/diagnostics.py) are the in-repo reference: all five are read-only, includingaf_link_identity, which returns a portal link but performs no mutation itself. -
Set
isError: trueon a failed call, not just error-shaped text. A tool that reports failure by returning ordinary text (or a string a client has to pattern-match) is indistinguishable from success to anything that checks the wire result rather than reads the prose — the broker's own audit pipeline learned this the hard way (9034858): a call that "failed" only in its text was counted as a success until the fix, and calls that correctly setisErrorare now the only ones auditedoutcome=error/error_class=tool_reported. Raising your SDK's tool-error exception (or building theCallToolResultexplicitly) gets you this automatically; returning a plain string never does. Expect this to move error rates in dashboards the day a backend ships it — that's the fix working, not a regression. -
Set
agent_policy— the model-facing per-service policy knob. The five conventions above are per-tool text the backend advertises;agent_policyis a per-service field you set in the Step 1 service list (alongsiderequired_permission/trust_tier) and is the model-facing half of dual enforcement (see docs/architecture.md's "Dual enforcement"). The aggregator composes every service'sagent_policyinto the MCP serverinstructionsthe LLM agent reads, so write 1–3 sentences of guidance the agent should reason over before calling this service's tools — which operations are safe reads, and which change real facility state and warrant confirming with the user first. Keep it distinct fromdescription(user-facing catalog UX shown in the portal):agent_policyaddresses the agent, imperatively. It is guidance, not an access-control boundary — therequired_permissiongate (maniaclab/af-mcp-platform#253also addstrust_tieras declared posture) remains authoritative. The shippedservices.yamlcarries one for each reference service — e.g. Rucio's says read queries are safe but creating or deleting rules changes real data placement and should be confirmed first.
Verification¶
# Check the broker sees the new service as a catalog entry
kubectl exec -n af-mcp deploy/af-mcp-broker -- \
curl -s http://localhost:8080/v1/catalog | jq '.servers[].name'
# Confirm the new service's entry has the fields you expect
kubectl exec -n af-mcp deploy/af-mcp-broker -- \
curl -s http://localhost:8080/v1/catalog | jq '.servers[] | select(.name=="my-new-service")'
/v1/catalog reports one entry per service, not per tool — per-tool
enumeration lives at GET /v1/catalog/{service}/tools (namespaced the same
way /mcp namespaces them). To confirm the actual tool names callers will
see end-to-end, talk to the aggregator's MCP protocol surface directly:
read -s -p "Bearer token: " MCP_BEARER_TOKEN
export MCP_BEARER_TOKEN
pixi run -e dev python scripts/verify-mcp-flow.py
This is the same script Connecting a Client recommends for sanity-checking a token — it prints every tool visible to the caller, grouped by inferred service prefix.