HTCondor credmon integration¶
How a job submitted through condor-mcp gets the user's grid credentials on the worker node -- an x509/VOMS proxy, a CERN Kerberos ticket, a ServiceX token -- without the LLM agent ever seeing one. The broker stays the only place credentials are minted; HTCondor's own credential pipeline (credd → credmon → shadow → starter) carries them to the job.
Code references: credmon/ (storer, sync loop, htcondor-api client, Vault
lease), app.py (the af-credmon/<kind> redeem audiences), api/admin.py
(/v1/admin/credmon), and the chart's broker.credmon values.
The chain¶
broker (every replica polls; the lease holder runs a cycle every interval)
for each user with a linked identity of kind K and a POSIX unixname:
mint top token = AF Broker Identity Token, aud=af-credmon/K, TTL 24h
POST htcondor-api /api/v1/creds/service/af_K {user, refresh: true}
→ credd writes <creddir>/<user>/af_K.top
AP credmon (one daemon, scans <creddir>/*/af_*.top)
POST broker /v1/credentials/K/redeem (Bearer: the top token)
→ writes <creddir>/<user>/af_K.use (ccache bytes / proxy PEM / token)
job: use_oauth_services = af_krb5, af_x509
shadow ships af_K.use → starter writes $_CONDOR_CREDS/af_K.use (0600,
the job's user), re-sent every SEC_CREDENTIAL_REFRESH (default 300s)
The broker keeps no copy of the top token: it is a signed, self-verifying
JWT, and credd holds the only copy. The kind is in the aud, so an
af-credmon/krb5 token is rejected by every other kind's redeem endpoint.
No condor service entry is involved -- each identity simply gains one more
consumer, resolved to that kind's first configured identity target (the same
default the user-facing /v1 surfaces use).
Revocation is by expiry plus the redeem endpoint's own live checks: a user who unlinks stops getting new top tokens, and their existing one stops redeeming as soon as the identity is gone. There is no per-token early revocation.
Broker configuration¶
Off by default. In the chart:
broker:
identityToken:
existingSigningKeySecret: broker-signing-key # top tokens are broker JWTs
credmon:
enabled: true
htcondorApiUrl: http://condor-mcp.mcp.svc.cluster.local:8080
existingHtcondorApiTokenSecret: credmon-htcondor-api-token # key: token
kinds: [krb5, x509] # empty = all three; unconfigured kinds skipped
# topTokenTtlSeconds: 86400 # must exceed syncIntervalSeconds
# syncIntervalSeconds: 14400
prometheusRule:
enabled: true
The equivalent env vars are CREDMON_ENABLED, CREDMON_HTCONDOR_API_URL,
CREDMON_HTCONDOR_API_TOKEN_FILE, CREDMON_SERVICE_PREFIX (default af_),
CREDMON_KINDS, CREDMON_TOP_TOKEN_TTL_SECONDS,
CREDMON_SYNC_INTERVAL_SECONDS and CREDMON_STATE_KV_PATH_PREFIX.
The broker refuses to boot when credmon is enabled without a signing key,
without Vault (users are enumerated from the Vault-backed identity stores,
and replicas coordinate through a Vault lease), without the Keycloak
principal directory (credd keys credentials by POSIX unixname), or with an
unreadable htcondor-api token file. A legacy-mode x509 entry (no
serviceUrl) has no Vault store, so x509 users are never enumerated there.
The htcondor-api bearer token must map to an identity in credd's
CRED_SUPER_USERS: the broker stores credentials on behalf of other users.
Access point prerequisites¶
On the AP (the schedd host), the site needs:
DAEMON_LIST = $(DAEMON_LIST) CREDD
SEC_CREDENTIAL_DIRECTORY_OAUTH = /var/lib/condor/oauth_credentials
CRED_SUPER_USERS = <identity of the htcondor-api bearer token>
plus the AP credmon daemon (see below). UChicago's head01 currently runs
DAEMON_LIST = MASTER SCHEDD -- no credd and no OAuth credential directory.
Coexisting with an existing condor_credmon_oauth. Which credmon writes a
.use file is decided by configuration, not timing. The Local, Client and
Pelican credmons only claim names listed in their *_PROVIDER_NAMES, and
the OAuth2 credmon only refreshes .use files that have a .meta file, so
none of them touch af_*. The Vault credmon's default * wildcard does
claim unclaimed .top files: a site running it must set an explicit
VAULT_CREDMON_PROVIDER_NAMES. credd keeps a single pid file and
CREDMON_COMPLETE per credential directory, so when another credmon
already owns them the AP credmon must run in its poll-only mode.
condor_submit checks credentials at submit time (CREDD_CHECK_CREDS): an
af_* service passes only if its .use file already exists. The storer
runs ahead of submission, so a linked user's files are normally in place; a
user who links an identity just before submitting may need to wait for the
next cycle, or an admin can run one now (below).
The AP credmon¶
The credmon is a long-running daemon on the AP (not part of this repository). Its contract:
- scan
<creddir>/*/<prefix>*.top; the kind is the filename minus the prefix and.top; - read
access_tokenfrom the.topJSON and callPOST <broker>/v1/credentials/<kind>/redeemwith it as the Bearer token -- the same contractaf_credentials.ProxyClientalready implements; - write
<creddir>/<user>/<prefix><kind>.useatomically (temp file in the same directory, mode 0400, root-owned, rename) before it expires; - on a 404 (identity no longer linked), remove the stale
.use; - primary mode: write
<creddir>/pid, rescan on SIGHUP, touchCREDMON_COMPLETEafter each scan. Poll-only mode: touch neither.
Jobs¶
#!/bin/bash
# Reference credentials by path only -- never print them.
export KRB5CCNAME="FILE:${_CONDOR_CREDS}/af_krb5.use"
export X509_USER_PROXY="${_CONDOR_CREDS}/af_x509.use"
klist -s && voms-proxy-info -exists
exec python analysis.py
Keep credentials out of job output. HTCondor does not redact job stdout
or stderr, and condor-mcp's output and exec tools return them to the agent
verbatim. Scripts must reference $_CONDOR_CREDS files only through
environment variables like the above -- never cat, base64 or echo
them. Only the short-lived .use credentials reach worker nodes; the top
token never leaves the AP credential directory. Keep
SEC_DEBUG_PRINT_KEYS off on the AP: it writes credentials into the
ShadowLog.
Operating it¶
The sync runs inside the broker on its own timer -- no CronJob or other scheduler to deploy. Every replica checks once a minute (the first check one minute after startup) whether a cycle is due; only the replica holding the per-cycle lease runs it, and every replica answers status queries identically from the shared Vault record.
- Portal → Admin: the HTCondor credmon panel shows the last cycle, its per-kind counts and errors, and a "Run sync now" button.
GET /v1/admin/credmon(admin only): the same record as JSON.POST /v1/admin/credmon/sync(admin only): run a cycle now from any replica; 409 while another replica is mid-cycle.- Metrics (metrics port):
af_mcp_credmon_sync_runs_total{outcome,trigger}(success,partial,failed,skipped_lease),af_mcp_credmon_sync_last_success_timestamp_seconds,af_mcp_credmon_sync_credentials_total{kind,outcome},af_mcp_credmon_sync_duration_seconds. Every replica publishes the shared last success, so alert ontime() - max(...) > 2 × interval-- the chart's optional PrometheusRule does exactly that. - Audit: one
credmon_top_token_storeline per user and kind, never the token itself.
A failing sync never fails /v1/readyz: an htcondor-api outage should not
take the gateway down. A partial cycle (some users failed) still counts as
a success for alerting; its errors are in the status record.
Known limitations¶
- Unlinking does not delete the user's
af_*entry from credd; the top token expires within its TTL and stops redeeming immediately. - A failed cycle is retried at the next interval, not sooner; the top-token TTL (default 6× the interval) covers missed cycles.