Skip to content

Route System One requests through KubeAI

KubeAI exposes POST /openai/v1/systemone and forwards it to POST /v1/systemone on a selected model server. This endpoint is an inference extension alongside chat, embeddings and reranking.

Supported engine values are OLlama, VLLM, LlamaCpp and SGLang. This integration assumes the selected server image already implements /v1/systemone. KubeAI adds routing and access to its existing inference lifecycle; it does not install an endpoint into the engine, add a wrapper, change the model server launcher or translate between engine payload formats.

Enable the feature on a Model

Install a KubeAI version and Model CRD containing the SystemOne feature. Upgrade the controller and CRD together using your normal installation method; an older CRD rejects SystemOne in spec.features.

For an existing Model, add SystemOne to its feature list, preserving the features it already serves:

spec:
  features: [TextGeneration, SystemOne]

Keep its existing engine, url, resource profile, image, arguments and credentials. The image must serve the endpoint at the model Pod's configured inference port. Declaring the feature enables KubeAI routing; it does not check backend endpoint availability or change startup/readiness probes.

The following example uses vLLM and a small Hugging Face model. Choose a model, image and resource profile suitable for your backend's System One implementation:

apiVersion: kubeai.org/v1
kind: Model
metadata:
  name: ticket-decisions
  labels:
    team: support
spec:
  engine: VLLM
  features: [TextGeneration, SystemOne]
  url: hf://Qwen/Qwen2.5-0.5B-Instruct
  resourceProfile: nvidia-gpu-l4:1
  minReplicas: 1
  loadBalancing:
    strategy: LeastLoad

The engine and URL combinations below illustrate how to use the other existing launchers. They are configuration examples, not a guarantee that the chart's default images or the example models support a particular decision schema.

Runtime spec.engine Example spec.url
Ollama OLlama ollama://qwen2.5:0.5b
vLLM VLLM hf://Qwen/Qwen2.5-0.5B-Instruct
llama.cpp LlamaCpp hf://Qwen/Qwen2.5-0.5B-Instruct-GGUF:Q4_K_M
SGLang SGLang hf://Qwen/Qwen2.5-0.5B-Instruct

Use the existing engine configuration for model formats, sources, hardware and credentials. Set spec.image when you need an explicitly tested server image; use a digest for reproducible deployments. Infinity and FasterWhisper are not accepted by this endpoint.

Discover and select models

List Models that advertise the feature:

curl --fail-with-body \
  'http://localhost:8000/openai/v1/models?feature=SystemOne'

Without a feature query, model discovery defaults to TextGeneration. In the request, model must be the Model's Kubernetes metadata.name, such as ticket-decisions. It is not a Hugging Face repository, an Ollama model tag or an SDK's default model name. KubeAI forwards this value unchanged; the server's existing model alias configuration must accept it.

The standard X-Label-Selector header restricts eligible Models. For example, X-Label-Selector: team=support selects the Model above. A Model not found under the supplied selectors returns 404.

Adapter selection through model_adapter or model:adapter is rejected. System One forwarding preserves the original envelope and does not perform the model-name rewriting used by other endpoints for adapters.

Send a JSON request

KubeAI requires a JSON object with a non-empty string model. All other fields are backend-defined. The example below illustrates a decision payload; confirm its field names and question types against your selected engine implementation.

curl --fail-with-body http://localhost:8000/openai/v1/systemone \
  -H 'Content-Type: application/json' \
  -H 'X-Label-Selector: team=support' \
  --data-binary '{
    "model": "ticket-decisions",
    "state": {"ticket": "The checkout is unavailable."},
    "questions": {
      "urgent": {"type": "noul", "instructions": "Does this need immediate attention?"},
      "team": {
        "type": "choice",
        "instructions": "Which team should investigate?",
        "criteria": {"billing": "Payment issues", "engineering": "Software failures"}
      }
    }
  }'

KubeAI validates the routing envelope, then forwards the original body bytes. It preserves JSON key order, unknown fields, numeric representations and engine-specific options. It does not validate questions, calculate decisions, normalize probabilities or convert schemas between engines. Missing Content-Type defaults to JSON; application/json; charset=utf-8 is accepted.

Backend status codes, response bodies and end-to-end response headers pass through the shared proxy. Backend validation failures therefore retain their own response format rather than becoming KubeAI routing errors.

Send multipart input

Use multipart/form-data only when the selected backend implements that format. KubeAI expects exactly one text form field named request, containing the JSON routing envelope. Put model inside that JSON field. A separate model form field does not select the destination.

curl --fail-with-body http://localhost:8000/openai/v1/systemone \
  -F 'request={"model":"ticket-decisions","state":"Inspect this image","questions":{"damaged":{"type":"noul","instructions":"Is the item damaged?"}}}' \
  -F 'image=@photo.png;type=image/png'

The remaining field names, file types and payload schema belong to the backend. KubeAI preserves the multipart boundary, headers, binary file contents and part order. The request field must be a form value, not an uploaded JSON file; missing or repeated request fields are rejected before contacting the engine.

Routing, limits and retry behavior

Requests use the existing model lookup, load balancing, scale-from-zero, active-request metrics and cancellation lifecycle. A valid request scales the selected Model to at least one replica and waits for an available address. Rejected requests do not trigger scaling. Models can retain minReplicas: 0 when cold-start latency is acceptable.

Use loadBalancing.strategy: LeastLoad for these workloads. KubeAI does not extract a text prefix from System One payloads; PrefixHash therefore has no payload-derived prefix to distribute them.

The complete body is limited to 32 MiB (33,554,432 bytes), including multipart boundaries, form values and files. KubeAI buffers the body to inspect the model and replay it during retries. A chunked incoming request is forwarded with the buffered body's known Content-Length.

The shared proxy's retry settings apply, including its default retry status codes 500, 502, 503 and 504. A retry can repeat inference after a backend has started work. This endpoint provides no exactly-once or deduplication guarantee. Set client deadlines to accommodate cold starts, and coordinate client retry policies with the proxy's retry configuration.

Condition HTTP status Behavior
Method other than POST 405 Allow: POST; no backend request
Invalid JSON or multipart routing envelope 400 No model scaling
Missing, blank or non-string model 400 No model scaling
Unsupported content type, engine, adapter or missing feature 400 No model scaling
Model not found, including selector mismatch 404 No model scaling
Body exceeds 32 MiB 413 No model lookup or scaling
Backend returns an error Backend status Response passes through, subject to retry settings
Backend does not expose /v1/systemone Backend status, commonly 404 No fallback to another inference API
Address acquisition deadline expires 504 Shared proxy releases request accounting

KubeAI-generated proxy errors use an error field; internal 5xx error details are logged rather than returned to callers. Client cancellation uses the shared request context and releases active-request and load-balancer accounting.

Verify a deployment

Before serving production traffic, check model discovery, JSON forwarding and one backend-defined decision request against each image you deploy. Check multipart input if the backend supports it, validation failures, an unavailable route, scale-from-zero and a client timeout. Readiness of the model server alone does not establish System One endpoint compatibility.

Use the normal inference ingress controls for authentication, TLS, timeouts, request size and concurrency. This route does not add a separate authentication service. Budget memory for concurrent buffered uploads and configure the ingress body limit consistently with the proxy limit.

Automated tests cover the four accepted engines using HTTP mock servers, including JSON/multipart byte preservation, chunked requests, retries, backend errors, feature gating and request accounting. CRD admission tests cover the new feature on all four engine values. These checks validate KubeAI integration; they do not establish compatibility with a particular real engine image or GPU.