Provisioner Configuration Reference
This document provides an exhaustive reference for all configuration parameters available in the Fuzzball central configuration system. The central configuration uses YAML format and supports node provisioners across multiple node provisioner backends with their specific parameters.
# Global cluster settings
nodeAnnotations:
# Map of global annotations applied to all nodes
# For example:
global.annotation: "cluster-wide-value"
environment: "production"
softwareTokens:
# Map of software license token limits
# For example:
matlab: 20
ansys: 10
scheduler:
queueDepth: 64
# Annotation keys that scheduler annotation matching should skip when
# comparing workflow job annotations against provisioner definitions'
# Resource.Annotations. Add keys that are routed by a provisioner-
# definition `policy:` expression rather than by Resource.Annotations.
ignoredAnnotations:
- nodepool
nodeHealth:
# Penalties and thresholds for the per-node reliability score.
# Omit entirely to accept the defaults.
cordonBelow: 65
image:
# Controls image cache write-back behavior.
# Omit entirely to accept the default (write-back enabled).
cacheWriteBack: true
nodeEventWebhooks:
# Endpoints that receive node health events as signed CloudEvents.
- url: https://alerts.example.com/fuzzball
secret: <shared secret>
definitions:
# Array of node provisioner definitions
# For example:
- id: compute-nodes
provisioner: static
# and more provisioner-specific configuration ...
priceAdjustments:
# Per-organization markups and discounts over the cluster's list rates
# For example:
- organization: 3f2b8c1e-9d4a-4f21-8f0e-2c7b6a1d5e93
multiplier: 1.2
Per-organization markups and discounts applied to node provisioner prices. The effective hourly price of a definition is its list rate multiplied by the organization’s multiplier, and that one number is used everywhere: scheduler placement, workflow cost estimates, the prices the organization sees when it lists node provisioners, and its charges.
An organization with no entry pays the list rate.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
organization | string | Yes | Organization UUID the adjustment applies to | 3f2b8c1e-9d4a-4f21-8f0e-2c7b6a1d5e93 |
multiplier | float | Yes | Scales the list rate of every definition for this organization. Must be greater than 0 and no greater than 100 | 1.2 |
definitions | map[string]float | No | Overrides multiplier for individual definitions, keyed by definition ID | aws-g5.xlarge: 0.9 |
Example:
priceAdjustments:
# A 20% markup on everything, with a negotiated discount on one GPU type
- organization: 3f2b8c1e-9d4a-4f21-8f0e-2c7b6a1d5e93
multiplier: 1.2
definitions:
aws-g5.xlarge: 0.9
# A flat 15% discount
- organization: 8c41f0aa-2b57-4c93-9e18-6d0a4f2b7c31
multiplier: 0.85
Keys under definitions are the definition IDs the cluster actually runs, after
AWS instance type expansion. A definition
declared as aws-${spec.instanceType} with instanceType: "g5.*" produces IDs
such as aws-g5.xlarge, and that expanded form is what belongs here – the
authored template never appears as a key.
List the real IDs with:
fuzzball node provisioner list
A key that names no configured definition is rejected when the configuration is set, so a typo fails loudly instead of quietly billing that definition at the list rate.
Because IDs are per definition rather than per instance type, a single instance
type can span several of them where a cluster declares -spot or -gpu
variants alongside the plain form. Each needs its own entry; the organization’s
multiplier covers everything not named.
A multiplier applies to workflows that start after it is set. Charges already recorded keep the rate they were placed at, so changing a markup never rewrites an invoice that has already been issued.
Removing adjustments follows the shape of the configuration document:
- Omitting
priceAdjustmentsentirely leaves the stored adjustments unchanged. - Setting
priceAdjustments: []clears every adjustment. - Listing some organizations updates those and leaves organizations the document does not mention alone.
Omitting the field is deliberately not the same as clearing it. A misspelledpriceAdjustmentskey is silently ignored when the configuration is parsed, and if absence meant “remove”, one typo would drop every organization back to list prices with no error. Clear adjustments with an explicit empty list.
Prices are reported to each organization as that organization pays them,
including the --max-cost-per-hour filter and --order cost sort on
fuzzball node provisioner list.
Cluster admins can read any organization’s effective prices, and additionally see the cluster’s pre-adjustment list rate:
fuzzball node provisioner list --organization 3f2b8c1e-9d4a-4f21-8f0e-2c7b6a1d5e93
Members of an organization always see their own prices, whatever they pass to
--organization, and never see the list rate behind them.
Price adjustments apply to compute. Storage, egress, and object cache are charged at their configured rates.
Global annotations applied to all cluster nodes.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
nodeAnnotations | map[string]string | No | Key-value pairs of annotations applied globally to all nodes | cluster.name: "production" |
Example:
nodeAnnotations:
cluster.name: "hpc-cluster-01"
datacenter: "us-west-2"
environment: "production"
cost.center: "research"
Software license token limits for concurrent usage control.
Software tokens are currently on the roadmap but not yet implemented.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
softwareTokens | map[string]uint32 | No | Software name to maximum concurrent license count mapping | matlab: 25 |
Example:
softwareTokens:
matlab: 25
ansys: 15
comsol: 8
abaqus: 10
scheduler:
# maximum number of requests in queue processed by scheduling iteration
queueDepth: 64
# how often the scheduler processes the queue
interval: 60s
# Cluster-wide expression (expr-lang) computing each allocation's scheduling
# priority every tick; a higher value is scheduled sooner. When unset, the
# default sums the organization, account, user, and workflow priority inputs
# and ages each allocation by one unit per hour spent in the queue.
priority: "organization.priority + account.priority + user.priority + workflow.priority"
# Preemption: allow a blocked higher-priority allocation to evict a
# preemptible, lower-priority running allocation. Disabled by default.
# Evictions only happen when they make the blocked allocation placeable.
preemptionEnabled: false
# Pool utilization percentage (0-100) at or above which preemption may
# evict; below it the blocked allocation is served by free or newly
# provisioned capacity instead.
preemptionThreshold: 75
preemptionGap: 10
minPreemptionRuntime: 60s
# Allow evicting a set of victims for one blocked allocation in a single
# scheduler pass (instead of at most one). Disabled by default.
preemptionMultiVictimEnabled: false
maxEvictionsPerTick: 8
preemptionDrainTimeout: 30s
# One-level (EASY-style) backfill: let lower-priority allocations fill a gap
# behind a blocked head allocation without delaying it. Enabled by default;
# set to false to schedule each node pool strictly in priority order — a
# blocked allocation then stops lower-priority work on the same pool for
# that pass.
backfillEnabled: true
# How long a running internal allocation (image or data fetch) is expected
# to hold its resources, used by the backfill availability estimate in
# place of the internal job's TTL. Estimate only: it never terminates a
# fetch that runs longer.
internalJobReleaseEstimate: 10m
# Federate deployments only: when auto-routing a workflow submission,
# discount each orchestrate cluster's score by how much of its ready queue
# outranks the submission on that cluster's own priority scale. Disabled by
# default; routing then picks purely by fit score and data locality. Set in
# the FEDERATE cluster's central config.
federationPriorityRoutingEnabled: false
# Blocking ratio (outranking ready allocations per usable node) at which a
# submission's routing score for a cluster halves.
federationRoutingCongestionThreshold: 1.0
# Annotation keys that scheduler annotation matching should skip when
# comparing job annotations against each candidate definition's
# Resource.Annotations. See "Scheduler annotation matching" below.
ignoredAnnotations:
- nodepool
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
queueDepth | uint32 | No | Scheduler queue depth (default: 64) | 128 |
interval | duration | No | How often the scheduler processes the queue (default: 60s). | 30s |
priority | string | No | Cluster-wide expr-lang expression that computes each allocation’s scheduling priority every tick (higher is scheduled sooner). When empty, defaults to organization.priority + account.priority + user.priority + workflow.priority plus one priority unit per hour the allocation has spent in the queue. The per-entity priority inputs are set by admins via group/organization/user update --priority; the per-workflow input via workflow start --priority (signed, <= 0). | "user.priority + workflow.priority" |
preemptionEnabled | bool | No | Enables the preemption pass, which may evict a preemptible, lower-priority running allocation in favor of a blocked higher-priority one (default: false). | true |
preemptionThreshold | float | No | Pool utilization percentage (0–100) at or above which the preemption pass may evict; 0 is treated as unset (default: 75). Utilization is measured per blocked allocation against the nodes it could actually run on, as the highest-utilized resource dimension it consumes (cores, memory, or a requested device kind such as GPUs — a saturated device kind the allocation does not request is ignored). Below the threshold — or while a dynamic provisioner definition can still provision nodes under its maxNodes cap — preemption is skipped and the blocked allocation waits for free or newly provisioned capacity instead. Note that utilization is a pool-level aggregate with no per-node fit awareness: a below-threshold pool whose free capacity is fragmented across nodes too small for the blocked allocation’s shape waits for natural drain rather than triggering preemption — a rising below_threshold miss count alongside a low pool_utilization gauge is the signal to lower the threshold. | 90 |
preemptionGap | float | No | Minimum effective-priority delta between a blocked allocation and an eviction candidate before that candidate may be preempted (default: 10). | 20 |
minPreemptionRuntime | duration | No | Anti-thrash floor: a running allocation cannot be preempted until it has been running at least this long (default: 60s). | 5m |
preemptionMultiVictimEnabled | bool | No | Allows the preemption pass to evict a set of victims for one blocked allocation in a single scheduler pass when no single victim frees enough capacity (default: false — at most one victim per blocked allocation per pass). Evictions always require that they make the blocked allocation placeable. | true |
maxEvictionsPerTick | uint32 | No | Caps the total victims the preemption pass may evict in one scheduler pass, across all blocked allocations (default: 8). | 16 |
preemptionDrainTimeout | duration | No | How long the preemption pass waits for an eviction’s freed capacity to appear before it may select new victims for the same blocked allocation (default: 30s). | 1m |
backfillEnabled | bool | No | Enables one-level (EASY-style) backfill, letting lower-priority allocations fill a gap behind a blocked head allocation without delaying it (default: true). Set to false to disable backfill; a blocked allocation then stops lower-priority work on the same node pool for that scheduling pass. | false |
internalJobReleaseEstimate | duration | No | How long a running internal allocation (image or data fetch) is expected to hold its resources, used by the backfill availability estimate in place of the internal job’s TTL (default: 10m). An estimate only — a fetch running past it is treated as releasing its resources imminently and is never terminated. Raise it on deployments where image pulls or data staging routinely take longer (e.g. slow WAN links) to keep availability estimates realistic. | 30m |
federationPriorityRoutingEnabled | bool | No | Federate deployments only, set in the federate cluster’s central config. When enabled, auto-routing a workflow submission (one without an explicit --cluster-id) additionally discounts each orchestrate cluster’s score by the amount of ready work that would run before the submission there: the cluster evaluates its own scheduler.priority expression for the candidate (including the submission’s --priority and any admin-set organization/group/user priorities it resolves locally) and counts the ready allocations that outrank it. Higher-priority work therefore routes into busy clusters it would jump the queue of, while lower-priority work prefers clusters where it starts sooner. Disabled by default — routing then picks purely by fit score and data locality. Clusters running a version that predates this feature report no blocking count and are scored as if nothing outranks the submission, so enable it only after every orchestrate cluster is upgraded. An explicit --cluster-id always pins the submission regardless of this setting. | true |
federationRoutingCongestionThreshold | float | No | Blocking ratio — outranking ready allocations per usable node on an orchestrate cluster — at which a submission’s routing score for that cluster halves; values <= 0 are treated as unset (default: 1.0). Lower values make routing flee outranking backlog sooner; higher values let fit score and data locality dominate. | 2.5 |
ignoredAnnotations | []string | No | Annotation keys to skip during scheduler annotation matching. A specific enumerated set of platform-internal fuzzball.io/* keys (e.g. fuzzball.io/workflow.id, fuzzball.io/job.name) is always ignored automatically, as is the dynamically keyed fuzzball.io/connect/<service> namespace — but otherwise this is an allowlist, not a prefix match, so user-defined keys placed under fuzzball.io/ are NOT auto-exempt. This list is for additional keys handled by policy: expressions on your provisioner definitions. | ["nodepool"] |
A node provisioner has two scoring knobs that determine its fitness for a given workflow job:
definition.annotations— values matched key-by-key against the job’s annotations by scheduler annotation matching, with built-in matchers per key (string equality by default; the GPU dimensions use substring or numeric-minimum matchers; see Built-in matchers below).definition.policy— an Expr expression evaluated against the job’s request; returns a boolean that gates eligibility.
The two knobs are independent: scheduler annotation matching does not look
at definition.policy, and policy evaluation does not look at
definition.annotations. Either can route on the same annotation key; a
deployment is free to use one, the other, or both.
By default every annotation key on a workflow job must be matched by an
entry in definition.annotations on each candidate definition; otherwise
the candidate is rejected for that job. The platform’s own annotation
keys — a specific enumerated set of fuzzball.io/* keys including
workflow/job/account identifiers (fuzzball.io/workflow.id,
fuzzball.io/job.name, …) and the provisioner-definition pinning key —
are skipped automatically, as is the dynamically keyed
fuzzball.io/connect/<service> namespace (the client-side connect
command, which cannot be enumerated as fixed keys). Apart from that
namespace this is an allowlist, not a prefix match: the fuzzball.io/
prefix is reserved for the platform, and any user-defined key placed
under that prefix is not auto-exempt and will still need to be added
to scheduler.ignoredAnnotations. Cluster admins should declare only
deployment-specific keys (their own labels for routing, etc.) in that
list.
Add an annotation key to ignoredAnnotations when routing for that key is
handled by definition.policy rather than by definition.annotations.
Without the entry, scheduler annotation matching would additionally
require the key on every candidate definition’s annotations map,
redundant with what the policy already evaluates.
For example, a deployment whose node provisioners look like
definitions:
- id: pool-small
provisioner: pbs
policy: |-
request.job_annotations["nodepool"] in ["pbs-small", "small", ""]
# ... no definition.annotations["nodepool"] needed; the policy handles it
should set:
scheduler:
ignoredAnnotations:
- nodepool
so that a workflow with resource.annotations.nodepool: small reaches the
policy without being rejected by scheduler annotation matching first.
Scheduler annotation matching uses string equality (ExactMatch) by
default. The following GPU dimensions have built-in non-exact matchers:
| Annotation key | Matcher |
|---|---|
nvidia.com/gpu.arch | ExactMatch |
nvidia.com/gpu.model | SubstringMatch (case-insensitive) |
nvidia.com/gpu.family | ExactMatch |
nvidia.com/gpu.product | SubstringMatch (case-insensitive) |
nvidia.com/gpu.memory | MinimumMatch (definition value ≥ requested) |
nvidia.com/gpu.compute.major | MinimumMatch |
nvidia.com/gpu.compute.minor | ExactMatch |
nvidia.com/gpu.count | MinimumMatch |
Fuzzball derives a reliability score from each node’s health conditions and the job failures
attributed to it — see
Node Health Monitoring.
The nodeHealth section tunes how that score is calculated. Omit it entirely to accept the
defaults.
The score is shown by
fuzzball node list,fuzzball node show, the API, the web UI and Prometheus. Whether automation acts on it is set bymode, which defaults toobserve— Fuzzball records what it would have done and takes no action. Raisemodeto let automation cordon, drain, or replace.
drainBelowhas one effect that does not depend onmode: the scheduler will not place new work on a node scoring below it, in any mode. Work already running there is left alone. Tuning this value therefore changes scheduling even on a cluster that has never enabled a policy — see Nodes below the drain threshold.
nodeHealth:
# Points deducted while each condition is active.
penalties:
MEMORY_ERRORS: 45
MACHINE_CHECK_ERRORS: 45
DISK_DEGRADED: 40
THERMAL_THROTTLE: 30
AGENT_UNHEALTHY: 15
# Score thresholds for each rung.
degradedBelow: 95
cordonBelow: 65
drainBelow: 45
# Attributed job failures.
failureTallyPenalty: 5
failureTallyCap: 25
failureTallyHalfLifeHours: 168
# How far automation may act on the score.
mode: observe
maxUnavailablePercent: 20
cordonSuppressionMinutes: 30
maxReplacementsPerHour: 3
# Fault classes that evacuate a node the moment they are raised.
evictOnFaultClasses:
- GPU_ECC_ERRORS
maxEvictionRestarts: 3
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
penalties | map | No | Points deducted from a starting score of 100 while a condition is active, keyed by condition name. Unlisted conditions keep the default penalty below | MEMORY_ERRORS: 45 |
degradedBelow | integer | No | Score below which a node is marked degraded (default: 95) | 95 |
cordonBelow | integer | No | Score below which automation cordons the node (default: 65) | 65 |
drainBelow | integer | No | Score below which automation drains the node, and below which no new work is placed on it (default: 45) | 45 |
failureTallyPenalty | integer | No | Points deducted per job failure attributed to the node (default: 5) | 5 |
failureTallyCap | integer | No | Ceiling on the total penalty from attributed failures, so a long history cannot condemn a node on its own (default: 25) | 25 |
failureTallyHalfLifeHours | integer | No | Half-life of the attributed-failure deduction, measured from the node’s most recent failure – so a new failure returns the whole tally to full weight (default: 168, one week) | 168 |
mode | string | No | How far automation may act: observe, cordon, drain or replace (default: observe) | observe |
maxUnavailablePercent | integer | No | Most of a definition’s nodes automation may hold out of service at once; 100 disables the ceiling (default: 20) | 20 |
cordonSuppressionMinutes | integer | No | How long automation leaves a node alone after a manual uncordon (default: 30) | 30 |
maxReplacementsPerHour | integer | No | How many of a definition’s cloud nodes replace mode may terminate in a rolling hour (default: 3) | 3 |
evictOnFaultClasses | list | No | Condition names that evacuate a node as soon as they are raised. Empty, the default, means nothing evicts | [GPU_ECC_ERRORS] |
maxEvictionRestarts | integer | No | How many times fault eviction may move one job before failing it instead (default: 3) | 3 |
Penalties apply only while a condition is active, so clearing one restores its points on the node’s next health report. Attributed failures decay instead of clearing. Both paths, and the way an uncordon deliberately does not reset the score, are described in How the score recovers.
mode is observe unless you change it, which means Fuzzball records what it
would have done and takes no action. See
Automated Response
for what each mode does and why observe is the default.
A node provisioner definition’s mode replaces the cluster’s rather than
merging with it, so a GPU pool can enforce while the rest of the cluster
observes.
An unrecognisedmodeis rejected when the configuration is validated rather than being treated asobserve. A typo here is most likely someone trying to enable enforcement, and silently leaving it off would look identical to the policy simply never firing.
| Condition | Default (critical) | Default (warning) | Detects |
|---|---|---|---|
MEMORY_ERRORS | 45 | 10 | Uncorrectable ECC errors; correctable errors climbing |
MACHINE_CHECK_ERRORS | 45 | 10 | Fatal machine checks; non-fatal machine check records |
DISK_DEGRADED | 40 | 10 | Failing SMART health or a read-only device; device I/O errors |
THERMAL_THROTTLE | 30 | 30 | A thermal zone at or above its critical trip point |
AGENT_UNHEALTHY | 15 | 15 | The local substrate runtime is not ready |
GPU_ECC_ERRORS | 45 | 10 | Uncorrectable GPU ECC errors; correctable errors climbing |
GPU_XID_ERRORS | 45 | 10 | Xid faults naming the card itself; Xids raised by a job’s own kernel |
GPU_THERMAL_THROTTLE | 30 | 30 | GPU clocks held down for a thermal or power reason |
COLLECTOR_FAILED | 0 | 0 | A collector ran and could not produce a reading |
A condition’s penalty depends on its severity as well as its name. Correctable memory errors
climbing and a single uncorrectable error are both MEMORY_ERRORS, but the first is a warning and
the second is critical. A penalties entry overrides both severities for that condition.
COLLECTOR_FAILED defaults to zero: a broken sensor is missing information rather than fault
evidence, so it is recorded on the node but does not move the score by default. Give it a penalty
if you would rather a node with unreadable sensors score lower.
A penalty here lowers the score but does not make automation act on the node. A node with a failed
collector is left alone whatever its score, exactly like a node whose telemetry has gone stale — see
Guardrails. No policy
action is taken and no policy event is emitted; the node is held at Unknown with the condition
attached, for an operator to judge.
HEARTBEAT_MISSEDandTELEMETRY_STALEnever contribute to the score. Apenaltiesentry for either is rejected when the configuration is validated, rather than accepted and then ignored. A node Fuzzball cannot see is recorded as unknown rather than scored as broken — otherwise a network interruption would look like failing hardware and could drain healthy capacity.
A node provisioner definition may set its own nodeHealth block, which replaces the cluster-wide
one for the nodes that definition provisions. Use it where a pool’s hardware or risk tolerance
differs from the rest of the cluster:
nodeHealth:
cordonBelow: 65
definitions:
- id: gpu-pool
provisioner: aws
# This pool reacts sooner than the rest of the cluster.
nodeHealth:
cordonBelow: 80
drainBelow: 60
provisionerSpec:
instanceType: p5.48xlarge
Resolution is per block, not per field: a definition that sets any nodeHealth value owns scoring
for its nodes outright and inherits nothing from the cluster block. Within a block, individual
fields still fall back to their defaults, so a definition can override one threshold without
restating the rest.
Controls how the workflow pipeline handles container image caching. Omit this section entirely to accept the defaults.
image:
# Whether to upload converted SIF images to the object cache after a
# cache-miss pull. Enabled by default.
cacheWriteBack: true
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
cacheWriteBack | bool | No | Controls whether the substrate uploads converted SIF images to the object cache after a cache-miss OCI→SIF conversion (default: true). When enabled, the first pull of an OCI image converts it to SIF and uploads the result to the cluster’s object cache; subsequent pulls of the same image on any node are served from the cache without re-converting. Set to false to disable the upload — the conversion still happens locally on each node, but the converted SIF is not uploaded back to the cache. Cache reads are unaffected: images already in the cache are still served from it regardless of this setting. | false |
Disable write-back on deployments where the network link between substrate nodes and the object cache is bandwidth-constrained — for example, segmented deployments where GPU nodes connect to the orchestrate cluster over a Tailscale or WireGuard tunnel. On such links, uploading a large SIF (8+ GB for ML framework images) after conversion can take hours and block the image stage, delaying job start.
With write-back disabled, each node converts OCI images locally on every cache miss. This trades redundant conversion work for faster job start times on slow links. On deployments with fast links between nodes and the object cache, leave write-back enabled (the default) — the one-time upload cost is repaid on every subsequent pull.
Disabling write-back does not affect images already in the object cache. A previously cached image is still served from the cache on a cache hit, regardless of this setting. Only the upload-after-conversion step on cache misses is skipped.
Federate deployments: the object cache feeds the data-locality bonus used by federate workflow routing. When write-back is disabled, newly pulled images do not create cache refs, so they stop contributing to the cluster’s locality score. On a federate deployment, disabling write-back on an orchestrate cluster reduces its attractiveness for workflows that reference images it has converted but not cached. If data locality matters for routing decisions, weigh this against the bandwidth savings before disabling write-back.
Example — segmented deployment with remote GPU nodes:
image:
cacheWriteBack: false
Endpoints that receive node health events as signed CloudEvents. See Webhook Notifications for the payload shape and how to verify a signature.
nodeEventWebhooks:
- url: https://alerts.example.com/fuzzball
secret: <shared secret>
events:
- POLICY_ACTED
- POLICY_HALTED
- NODE_REPLACED
timeoutSeconds: 5
maxAttempts: 3
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
url | string | Yes | Endpoint events are POSTed to. Must be http or https | https://alerts.example.com/fuzzball |
secret | string | Yes | Shared secret used to sign every request with HMAC-SHA256 | <shared secret> |
events | list | No | NodeEvent type names to deliver. Empty, the default, delivers every event | [POLICY_ACTED] |
timeoutSeconds | integer | No | Bound on a single delivery attempt (default: 5) | 5 |
maxAttempts | integer | No | Attempts before an event is dropped (default: 3) | 3 |
Unlike nodeHealth, this is cluster-scoped only. A definition’s nodeHealth block replaces the
cluster’s outright, so webhooks living inside it would be silently dropped by any definition that
tuned a single penalty — and losing notifications by accident is the failure this exists to
prevent.
secretis required. Every payload names a node and the fault taking it out of service, and a receiver with no way to tell a real delivery from a forged one cannot safely act on that. An endpoint configured without one is rejected when the configuration is validated.
Bounds where cluster members may point user-created hostpath and NFS storage provisioners. When
enabled, the privileged drivers (hostpath, nfs) become cluster-admin only, and any allowlists
below constrain where those admins may still point them. Multi-tenant deployments turn this on
automatically; single-tenant deployments leave it off unless an operator opts in for the same
guardrails. Distinct from defaultStorage, which describes the provisioner Fuzzball creates
itself — see the
storage configuration guide.
storageRestrictions:
enabled: true
hostPathRoots:
- /mnt/fuzzball
nfsServers:
- nfs-prod.internal
- nfs-scratch.internal
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
enabled | boolean | No | Restrict privileged driver creation (hostpath, nfs) to cluster administrators. On automatically in SaaS mode; can be set explicitly on any deployment (default: false) | true |
hostPathRoots | list | No | Absolute-path prefixes that hostpath provisioners may live under. Empty accepts any absolute path | ["/mnt/fuzzball"] |
nfsServers | list | No | Hostnames or addresses that NFS provisioners may target. Empty accepts any host | ["nfs.example.com"] |
The zero value imposes no restrictions. On operator-managed Kubernetes deployments, the operator
also populates hostPathRoots with the shared hostpath location it configures for the cluster.
definitions:
- id: p5.48xlarge
provisioner: aws
# Automatically run hardware discovery for this definition after a
# configuration update. Disabled by default; discovery can always be
# triggered manually with `fuzzball node provisioner discover`.
autoDiscover: true
provisionerSpec:
instanceType: p5.48xlarge
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
autoDiscover | bool | No | Runs hardware discovery for this definition automatically after each configuration update (default: false; not valid for static definitions). Definitions for instance types without node-reported or catalog data start with a generated approximation of their hardware; discovery boots one instance of the type, records the node’s hardware report — real CPU topology, usable memory, and full device details, including annotations that hardware-targeting workflows match against — and deletes the instance. Discovery instances are billable cloud instances, which is why automatic discovery is opt-in per definition; fuzzball node provisioner discover triggers the same process manually regardless of this setting. | true |
Each entry in the definitions array is a
node provisioner: a configuration that,
for a chosen backend, tells Fuzzball how to obtain a class of functionally-identical compute nodes
with policy attached. The serialized form of a node provisioner is referred to as a node provisioner
definition.
These parameters are available for node provisioners across all backends:
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
id | string | Yes | Unique identifier for the provisioner definition | "compute-nodes" |
annotations | map[string]string | No | Key-value pairs of annotations specific to this definition | node.type: "compute" |
provisioner | string | Yes | Node provisioner backend: static, aws, gcp, azure, slurm, pbs, coreweave, oci | "static" |
policy | string | No | Expression-based policy controlling access to this definition | request.owner.organization_id == "research" |
ttl | uint32 | No | Node lifetime in seconds after provisioning. Required and must be > 0 for pbs and slurm definitions; must be 0 or omitted for static definitions (a non-zero value is rejected). | 86400 |
ttlBuffer | uint32 | No | Per-node buffer added to the allocation TTL, scaled by the number of nodes in the allocation. Accounts for provisioning and scheduling delays in multi-node jobs. Must be 0 or omitted for static definitions. Ignored when 0. | 300 |
exclusive | string | No | Node exclusive level: empty or none (default, shared), job (exclusive to one job), or workflow (exclusive to one workflow) | "job" |
maxNodes | uint32 | No | Maximum size of this definition’s node pool (dynamic definitions only, clamped to the 128 backstop). See Node Pool Capping with maxNodes | 64 |
nodeHealth | object | No | Reliability scoring for this definition’s nodes, replacing the cluster-wide block. See Per-definition scoring | cordonBelow: 80 |
registrationDeadline | duration | No | How long the scheduler waits for the nodes in a provisioning request to register before it deletes the instances that request created and fails the workflow (default: 15m). Ignored for static definitions. See Node Registration Deadline | 20m |
provisionerSpec | object | Yes | Provisioner-specific configuration (see sections below) | - |
ttlandttlBufferare bothuint32values. When a non-static provisioner definition hasttlBuffer > 0and a non-zero allocation TTL is being set for a node, the scheduler addsttlBuffer × nodeCountto the allocation TTL before submitting the provisioning request. Both the multiplication and addition use saturating arithmetic: if either result would exceed 4,294,967,295 (roughly 136 years), it is clamped to that value rather than wrapping around. This prevents misconfigured large values from silently producing a much shorter TTL and causing nodes to be terminated before jobs complete.This formula does not apply to static provisioners, or when either the allocation TTL or
ttlBufferis 0.
The exclusive parameter controls how nodes provisioned by this definition are shared among jobs:
If not specified or empty, nodes are shared and can run multiple jobs simultaneously. Multiple jobs from the same or different workflows can be scheduled on the same node based on available resources.
job: Nodes are exclusive to a single job allocation. Once a job is assigned to the node, no other jobs can use it until the job completes and the node is cleaned up. This ensures complete isolation at the job level.workflow: Nodes are exclusive to a single workflow. All jobs within the same workflow can share the node, but jobs from other workflows cannot use it. This is useful for workflows that need dedicated resources but want to share nodes across their jobs.
Example:
definitions:
# Shared nodes for general workloads
- id: shared-compute
provisioner: static
exclusive: none
provisionerSpec:
condition: hostname() matches "shared-[0-9]+"
# Job-exclusive nodes for sensitive workloads
- id: exclusive-compute
provisioner: pbs
exclusive: job
ttl: 3600
provisionerSpec:
cpu: 8
memory: "32GiB"
queue: "workq"
The maxNodes parameter caps the size of a dynamic definition’s node pool: the number of nodes
provisioned for that definition at any one time.
- Pool cap: The pool is capped at
min(maxNodes, 128). The cap applies to dynamic definitions only; static definitions are not affected. - Blocking: When the pool is at the cap, an allocation that fits within the cap waits for an
existing node to become free (node reuse) instead of provisioning new nodes. The scheduler
publishes a
provisioning_blockedworkflow event (carryingdefinition_id,total_nodes, andmax_nodesattributes) and a genericscheduling_blockedevent, once per blocked allocation. See Troubleshooting scheduling_blocked events. - Rejection: A multi-node job that alone requires more nodes than the definition allows fails at workflow submission. Task arrays are not rejected; their concurrent width is clamped instead.
- Unset maxNodes: If
maxNodesis not configured, the pool can grow to the 128 backstop, but a single allocation is still limited to the default of 16 nodes. - Visibility: The definition’s
maxNodesvalue appears infuzzball node provisioner getoutput and as Max Nodes in the web UI’s provisioner details (a dynamic definition without an explicit value reports the default of 16; an uncapped static definition omits the field).
Example:
definitions:
# Pool capped at 32 nodes
- id: small-pool
provisioner: pbs
ttl: 3600 # node lifetime set to 1h
maxNodes: 32
provisionerSpec:
cpu: 8
memory: "32GiB"
queue: "workq"
# Pool capped at the 128 backstop (maxNodes not set)
- id: large-pool
provisioner: slurm
ttl: 7200 # node lifetime set to 2h
provisionerSpec:
cpu: 16
memory: "64GiB"
partition: "compute"
In the first example, once 32 small-pool nodes are provisioned, further allocations wait for a
node to free up. In the second example, the pool can grow up to 128 nodes before blocking.
Every dynamic backend provisions asynchronously. The provisioning request returns as soon as the
backend accepts it, long before the nodes have finished booting and installing the Fuzzball
Substrate. If that bootstrap then fails — a cloud-init error, a package install failure, an NFS
mount that never comes up — a node never registers with the scheduler, and nothing in the backend
reports a problem. registrationDeadline bounds how long the scheduler waits for those nodes
before it gives up.
- Clock start: The deadline runs from the moment the provisioning request returns, not from
when the job entered the Fuzzball queue. An allocation can sit in the queue for a long time
before it provisions (behind the definition’s
maxNodescap, higher-priority work ahead of it, or overall queue depth), and that wait does not count against the deadline. - Expiry behavior: The scheduler cordons any nodes that did register, deletes every instance created by the provisioning request, then fails the workflow with an error naming how many nodes registered out of how many were requested. Expiry is terminal: the workflow does not retry and must be resubmitted.
- Partial registration: A request for three nodes where only one node registered expires exactly as a request where none registered. Once every node in a request has registered, the deadline no longer applies to it — a node lost after that point is handled as a node failure, not as a registration timeout.
- Check frequency: The scheduler evaluates deadlines every 30 seconds, so expiry is detected shortly after the deadline passes.
- Static definitions: These never issue a provisioning request, so a value set on a static definition is ignored rather than rejected.
Onslurmandpbsdefinitions the provisioning request returns when the batch job is submitted withsbatchorqsub, not when it starts running. Time the batch job spends queued in Slurm or PBS therefore counts againstregistrationDeadline. On a busy backend where jobs routinely wait longer than the15mdefault, raiseregistrationDeadlineon those definitions — otherwise the scheduler fails the workflow of a legitimately queued batch job while that job is still waiting for backend resources. Cleanup signals the batch job rather than deleting it, and a job that has not started yet rejects the signal: it is left in the backend queue and must be removed withscancelorqdel.
Write the value as a duration with a unit suffix. An unsuffixed integer (900), a negative value
(-5m), or anything below 1s is rejected when the configuration is loaded, rather than clamped
or silently replaced with the default. Omit the field to use the 15m default.
Example:
definitions:
# Large GPU images in a slow region: allow longer than the 15m default.
- id: gpu-pool
provisioner: aws
ttl: 7200
registrationDeadline: 30m
provisionerSpec:
instanceType: p5.48xlarge
# A busy Slurm partition: the deadline must cover backend queue wait.
- id: batch-pool
provisioner: slurm
ttl: 14400
registrationDeadline: 2h
provisionerSpec:
cpu: 16
memory: "64GiB"
partition: "compute"
To diagnose a workflow that failed this way, see Troubleshooting provisioning_registration_timeout events.
Static provisioners manage physical or pre-allocated compute resources.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
condition | string | Yes | Expression-based condition for node matching | hostname() matches "compute-[0-9]+" |
costPerHour | float64 | No | Cost per hour for resource usage (must be ≥ 0) | 0.25 |
The condition field supports these built-in variables and functions:
| Variable | Type | Description | Example Value |
|---|---|---|---|
uname.sysname | string | Operating system name | "Linux" |
uname.nodename | string | Network node hostname | "compute-001" |
uname.release | string | Operating system release | "5.4.0-74-generic" |
uname.version | string | Operating system version | "#83-Ubuntu SMP" |
uname.machine | string | Hardware machine type | "x86_64", "aarch64" |
uname.domainname | string | Network domain name | "cluster.local" |
| Variable | Type | Description | Example Value |
|---|---|---|---|
osrelease.name | string | OS name | "Ubuntu" |
osrelease.id | string | OS identifier | "ubuntu" |
osrelease.id_like | string | Similar OS identifiers | "debian" |
osrelease.version | string | OS version string | "20.04.3 LTS (Focal Fossa)" |
osrelease.version_id | string | OS version identifier | "20.04" |
osrelease.version_codename | string | OS version codename | "focal" |
| Variable | Type | Description | Example Value |
|---|---|---|---|
cpuinfo.vendor_id | string | CPU vendor | "GenuineIntel", "AuthenticAMD" |
cpuinfo.cpu_family | uint | CPU family number | 6 |
cpuinfo.model | uint | CPU model number | 158 |
cpuinfo.model_name | string | CPU model name string | "Intel(R) Xeon(R) CPU E5-2680 v4" |
cpuinfo.microcode | uint | Microcode version | 240 |
cpuinfo.cpu_cores | uint | Number of physical CPU cores | 16 |
| Function | Return Type | Description | Example |
|---|---|---|---|
hostname() | string | Returns current hostname | "compute-001" |
modalias.match(pattern) | bool | Matches hardware modalias patterns | modalias.match("pci:v000010DEd*") |
# NVIDIA GPU (any model)
modalias.match("pci:v000010DEd*sv*sd*bc03sc*i*")
# Specific NVIDIA GPU models
modalias.match("pci:v000010DEd00001B06sv*sd*bc03sc*i*") # GTX 1080 Ti
modalias.match("pci:v000010DEd00001E07sv*sd*bc03sc*i*") # RTX 2080 Ti
# Intel Ethernet controllers
modalias.match("pci:v00008086d*sv*sd*bc02sc00i*")
# Mellanox InfiniBand adapters
modalias.match("pci:v000015B3d*sv*sd*bc0Csc06i*")
You can also easily get the modalias for all the PCI devices on a node to match a specific device with the following one-liner:
$ IFS=$'\n'; for d in $(lspci); do modalias=$(cat /sys/bus/pci/devices/0000\:${d%% *}/modalias); echo "$modalias -> ${d#* }"; done
pci:v00008086d00004641sv00001D05sd00001174bc06sc00i00 -> Host bridge: Intel Corporation 12th Gen Core Processor Host Bridge/DRAM Registers (rev 02)
pci:v00008086d0000460Dsv00000000sd00000000bc06sc04i00 -> PCI bridge: Intel Corporation 12th Gen Core Processor PCI Express x16 Controller #1 (rev 02)
pci:v00008086d000046A6sv00001D05sd00001174bc03sc00i00 -> VGA compatible controller: Intel Corporation Alder Lake-P GT2 [Iris Xe Graphics] (rev 0c)
[snip...]definitions:
# Basic compute nodes
- id: compute-standard
provisioner: static
provisionerSpec:
condition: |-
hostname() matches "compute-[0-9]{3}" &&
cpuinfo.vendor_id == "GenuineIntel" &&
cpuinfo.cpu_cores >= 16
costPerHour: 0.40
# GPU nodes
- id: gpu-nodes
provisioner: static
provisionerSpec:
condition: |-
hostname() matches "gpu-[0-9]+" &&
modalias.match("pci:v000010DEd*sv*sd*bc03sc*i*")
costPerHour: 2.50
# High-memory nodes
- id: highmem-nodes
provisioner: static
provisionerSpec:
condition: |-
hostname() matches "mem-[0-9]+" &&
cpuinfo.cpu_cores >= 64
costPerHour: 1.75
condition: |-
osrelease.id == "ubuntu" &&
osrelease.version_id >= "20.04"
condition: |-
uname.machine == "x86_64" &&
cpuinfo.vendor_id == "GenuineIntel" &&
cpuinfo.cpu_cores >= 16
condition: |-
let compute_regex = "compute-[0-9]{3}";
let gpu_regex = "gpu-[0-9]{2}";
hostname() matches compute_regex || hostname() matches gpu_regex
condition: |-
// Match NVIDIA GPU devices
modalias.match("pci:v000010DEd*sv*sd*bc03sc*i*") &&
cpuinfo.cpu_cores >= 8
condition: |-
let is_compute_node = hostname() matches "compute-[0-9]+";
let is_intel_cpu = cpuinfo.vendor_id == "GenuineIntel";
let is_ubuntu = osrelease.id == "ubuntu";
let has_enough_cores = cpuinfo.cpu_cores >= 16;
is_compute_node && is_intel_cpu && is_ubuntu && has_enough_cores
AWS provisioners support dynamic EC2 instance provisioning.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
instanceType | string | Yes | EC2 instance type or wildcard pattern | "t3.large", "c5.*" |
spot | bool | No | Use spot instances | true, false |
AWS provisioners support wildcard patterns that automatically expand to individual instance types:
t3.*expands tot3.nano,t3.micro,t3.small, etc.c5.*expands toc5.large,c5.xlarge,c5.2xlarge, etc.p3.*expands top3.2xlarge,p3.8xlarge,p3.16xlarge
When using wildcards, the ${spec.instanceType} placeholder in the definition ID is replaced with the actual instance type.
definitions:
# Spot instances for cost optimization
- id: aws-${spec.instanceType}-spot
provisioner: aws
provisionerSpec:
instanceType: t3.*
spot: true
policy: |-
request.job_ttl <= 3600
# On-demand compute instances
- id: aws-${spec.instanceType}
provisioner: aws
provisionerSpec:
instanceType: c5.*
spot: false
policy: |-
request.job_kind == "service"
# GPU instances for ML workloads
- id: aws-${spec.instanceType}-gpu
provisioner: aws
provisionerSpec:
instanceType: p3.*
spot: false
policy: |-
request.job_resource.devices["nvidia.com/gpu"] > 0
Slurm provisioners integrate with existing Slurm clusters.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
costPerHour | float64 | No | Cost per hour for resource usage (must be ≥ 0) | 0.30 |
cpu | int | Yes | Number of CPU cores (must be > 0) | 16 |
memory | string | Yes | Memory specification | "64GiB" |
partition | string | Yes | Slurm partition name | "compute" |
definitions:
# Standard compute partition
- id: slurm-compute
provisioner: slurm
ttl: 86400 # node lifetime set to 24h
provisionerSpec:
costPerHour: 0.30
cpu: 16
memory: "64GiB"
partition: "compute"
policy: |-
request.job_resource.cpu.cores <= 16
# GPU partition
- id: slurm-gpu
provisioner: slurm
ttl: 43200 # node lifetime set to 12h
provisionerSpec:
costPerHour: 1.80
cpu: 8
memory: "32GiB"
partition: "gpu"
policy: |-
request.job_resource.devices["nvidia.com/gpu"] > 0
PBS provisioners integrate with OpenPBS/PBS Pro clusters.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
cpu | int | Yes | Number of CPU cores (must be > 0) | 8 |
memory | string | Yes | Memory specification | "32GiB" |
gpus | int | No | Number of GPUs (must be ≥ 0) | 1 |
queue | string | Yes | PBS queue name | "workq" |
costPerHour | float64 | No | Cost per hour for resource usage (must be ≥ 0) | 0.30 |
definitions:
# Standard PBS queue
- id: pbs-compute
provisioner: pbs
ttl: 86400 # node lifetime set to 24h
provisionerSpec:
cpu: 8
memory: "32GiB"
gpus: 0
queue: "workq"
costPerHour: 0.30
# GPU queue
- id: pbs-gpu
provisioner: pbs
ttl: 86400 # node lifetime set to 24h
provisionerSpec:
cpu: 4
memory: "16GiB"
gpus: 1
queue: "gpu"
costPerHour: 1.00
CoreWeave provisioners support dynamic instance provisioning on CoreWeave’s cloud infrastructure, optimized for GPU workloads.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
instanceType | string | Yes | CoreWeave instance type identifier | "cd-a40-24gb" |
costPerHour | float64 | No | Cost per hour for resource usage (must be ≥ 0) | 15.00 |
definitions:
# CPU instance for standard compute workloads
- id: coreweave-cpu-small
provisioner: coreweave
provisionerSpec:
instanceType: "cd-gp-a192-genoa"
costPerHour: 7.78
policy: |-
request.job_resource.cpu.cores >= 4 &&
request.job_resource.cpu.cores <= 32
# GPU instance for ML/AI workloads
- id: coreweave-gpu-a40
provisioner: coreweave
ttl: 3600
provisionerSpec:
instanceType: "cd-a40-24gb"
costPerHour: 15.00
policy: |-
request.job_resource.devices["nvidia.com/gpu"] > 0
CoreWeave provisioners support both dynamic provisioning (on-demand node creation) and static provisioning (pre-existing node pools). For static provisioning with pre-existing CoreWeave node pools, see the CoreWeave Static Provisioning guide.
OCI provisioners support dynamic instance provisioning on Oracle Cloud Infrastructure.
| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
shape | string | Yes | OCI compute shape. Flex shapes must encode their size as shape:ocpus:memoryGB | "VM.Standard.E4.Flex:4:32", "VM.GPU.A10.1" |
definitions:
# CPU instance for standard compute workloads
- id: oci-cpu-standard
provisioner: oci
provisionerSpec:
shape: "VM.Standard.E4.Flex:8:64"
policy: |-
request.job_resource.cpu.cores >= 4 &&
request.job_resource.cpu.cores <= 8
# GPU instance
- id: oci-gpu-a10
provisioner: oci
ttl: 3600
provisionerSpec:
shape: "VM.GPU.A10.1"
policy: |-
request.job_resource.devices["nvidia.com/gpu"] > 0
The OCI pricing API does not publish per-GPU rates, so cost estimates for GPU shapes always use a fallback per-GPU hourly rate: the built-in default of $3.00/GPU/hr, or thedefaultGPUHourlyRatevalue from the orchestrator’s OCI provisioner configuration when set.
Policy expressions control access to node provisioners and use the same expression language as static conditions.
Policies apply to every allocation, including the internal jobs Fuzzball
creates implicitly for container image pulls and data staging
(request.job_kind == "internal"). Internal jobs have a safety net user jobs
do not: if every definition’s policy rejects an internal job, the scheduler
falls back to placing it as if no policies were configured — a policy can
never make image pulls or data staging unschedulable. A policy expression
that fails to evaluate for an internal job is treated as rejecting that
definition rather than failing the workflow; if no definition passes as a
result, the same fallback applies — and may still place the internal job on
the definition whose policy could not be evaluated. Each time the fallback
places an internal job, the orchestrator logs a warning naming the
allocation and the definition chosen — if internal jobs are meant to be
routed somewhere specific (for example, image pulls to a data transfer
node), this warning is the signal that no definition’s policy admits them
and the routing policies need attention.
| Variable | Type | Description | Example |
|---|---|---|---|
request.owner.id | string | User ID | "user-123" |
request.owner.organization_id | string | Organization ID | "org-research" |
request.owner.email | string | User email address | "user@example.com" |
request.owner.cluster_id | string | Cluster ID | "cluster-01" |
request.owner.account_id | string | Group ID | "account-456" |
| Variable | Type | Description | Example |
|---|---|---|---|
request.job_kind | string | Job type | "job", "service", "internal" |
request.job_ttl | int | Job time-to-live in seconds | 3600 |
request.job_annotations | map[string]string | Job annotation key-value pairs | request.job_annotations["tier"] |
request.multinode_job | bool | True for multi-node jobs | true |
request.task_array_job | bool | True for task array jobs | false |
request.multinode_nodes | int | Node count requested by a multi-node job (0 otherwise) | 4 |
request.task_array_concurrency | int | Concurrency requested by a task array job (0 otherwise) | 8 |
| Variable | Type | Description | Example |
|---|---|---|---|
definition.nodes | int | The size of the pool this definition can offer an allocation. For a static definition, the count of usable nodes as they exist — including fully allocated ones, since a busy pool is still a pool — capped by an explicitly configured maxNodes. For a dynamic definition, its maxNodes capacity (what it may grow to). | definition.nodes >= request.multinode_nodes |
definition.max_nodes | int | The definition’s maxNodes: the configured value when set; otherwise unlimited for a static definition, or the default cap (16) for a dynamic one. | definition.max_nodes >= 4 |
| Variable | Type | Description | Example |
|---|---|---|---|
request.job_resource.cpu.affinity | string | CPU affinity | "none", "core", "socket", "numa" |
request.job_resource.cpu.cores | int | Number of CPU cores requested | 4 |
request.job_resource.cpu.threads | bool | Hyperthreading enabled | true |
request.job_resource.cpu.sockets | int | Number of CPU sockets | 1 |
request.job_resource.mem.bytes | int | Memory in bytes | 4294967296 |
request.job_resource.mem.by_core | bool | Memory allocation per core | false |
request.job_resource.devices | map[string]uint32 | Device requests | request.job_resource.devices["nvidia.com/gpu"] |
request.job_resource.exclusive | bool | Exclusive node access | true |
policy: |-
request.owner.organization_id == "research" &&
request.owner.account_id in ["2f0a8f4e-0a16-47d5-b541-05d3f9f44910", "c602cf05-7604-4f11-a690-79552b1fdbdd"]
policy: |-
request.job_resource.cpu.cores <= 32 &&
request.job_resource.mem.bytes <= (256 * 1024 * 1024 * 1024) &&
!request.job_resource.exclusive
policy: |-
request.job_kind == "job" &&
request.job_ttl >= 300 &&
request.job_ttl <= 86400
policy: |-
let gpu_count = request.job_resource.devices["nvidia.com/gpu"];
gpu_count > 0 && gpu_count <= 4 &&
request.owner.organization_id == "280abb59-b765-4cdd-a538-6ab8f9b7927c"
Reject multi-node or task-array jobs that can never be satisfied by this
definition’s pool. Because definition.nodes counts a static pool’s usable
nodes even while they are fully occupied, a busy pool keeps accepting jobs
(they queue until nodes drain) instead of being rejected at its busiest.
definition.nodes never exceeds definition.max_nodes — an explicitly
configured maxNodes already caps it — so gating on definition.nodes
alone is sufficient.
policy: |-
request.multinode_nodes <= definition.nodes &&
request.task_array_concurrency <= definition.nodes
Route container image pulls and data staging to a dedicated transfer node and keep them off the compute nodes, so user workloads do not compete with network-heavy transfers:
definitions:
- id: dtn
provisioner: static
provisionerSpec:
condition: hostname() == "fz-dtn01"
policy: request.job_kind == "internal"
- id: compute
provisioner: static
provisionerSpec:
condition: hostname() matches "fz-cmp0[1-3]"
policy: request.job_kind != "internal"
A definition without a policy accepts every job kind, so the attract policy
on the transfer node definition is not enough by itself: every other
definition needs the repel policy (request.job_kind != "internal"), or
internal jobs remain eligible there and may still be placed on compute nodes
by resource and cost scoring.
Pulled images are shared per storage segment, so the transfer node must be in the same segment as the compute nodes that run the jobs — otherwise those jobs cannot find the pulled image. If the transfer node pool has no usable nodes (for example, the node is down), internal jobs wait for it like any policy-gated job; they do not spill onto the compute definitions.
policy: |-
request.job_annotations["priority"] == "high" &&
request.job_annotations["project"] in ["proj-a", "proj-b"] &&
request.owner.email matches "*@ciq.com"
policy: |-
request.multinode_job ?
request.job_resource.cpu.cores >= 4 &&
request.owner.account_id == "092403fe-12ef-4465-bce4-18292fec13c8"
:
request.job_resource.cpu.cores <= 16
policy: |-
let current_hour = time.Now().Hour();
let is_business_hours = current_hour >= 9 && current_hour <= 17;
request.job_annotations["priority"] == "low" ? !is_business_hours : true