Skip to main content

Configuration

Client::openai(url) returns a ClientBuilder. Every method takes self and returns it, and build() returns Result<Client, Error>: the URL, the key source, and the headers are checked there, not on the first request.

use std::time::Duration;

use svir::layer::{Retry, Timeout};
use svir::prelude::*;

fn client() -> Result<Client, Error> {
Client::openai("https://models.example.com/v1")
.api_key_env("MODELS_API_KEY")
.connect_timeout(Duration::from_secs(5))
.layer(Retry::connect(3))
.layer(Timeout::first_token(Duration::from_secs(120)))
.build()
}
MethodEffectDefault
api_key(key) / api_key_env(name) / api_key_file(path)Bearer authenticationNone
header(name, value)A header sent with every request, for a gateway or a hosted endpointNone
allow_http()Plain HTTP to a host that is not loopbackRefused
connect_timeout(d)How long connecting may take10 s
idle_timeout(d) / no_idle_timeout()How long the server may send nothing5 min
include_usage(bool)Ask for token counts when a request does not sayOn
lenient() / mode(Mode)How strictly responses are readStrict
limits(Limits)Bounds on one response64 MiB, 256 KiB per event, 64 tool calls
think(Think)What to do with inline <think> tagsSplit into reasoning
context_tokens(n)Refuse a request that cannot fit, before sending itOff
http(backend)Another HTTP backendhyper
layer(l) / wrap(closure)Middleware around every callNone

Build one Client per server and clone it. Clones are cheap and share the connection pool and what was learned about the server. A client's Debug shows the URL and never the key or a header's value.

Strictness and limits have a page of their own: Strict decoding and limits.

The base URL​

Pass the server's base, with or without /v1 and a trailing slash; svir appends /v1/chat/completions and /v1/models. A path in front is kept: http://host/openai/v1 calls http://host/openai/v1/chat/completions.

build() fails with ErrorKind::Config for:

The URLWhy
http:// to a host that is not loopbackKeys and prompts would cross the network in the clear. Use https://, or .allow_http() for a trusted network
https:// without the tls or the tls-aws-lc featureNeither is on; see Features and TLS
Credentials, a query, or a fragment in the URLA key belongs in api_key, not in the URL
No scheme (127.0.0.1:1234)Write http://127.0.0.1:1234

Loopback is localhost, 127.0.0.1, and [::1].

API keys​

SourceRead
.api_key("...")The value given
.api_key_env("NAME")The variable, when the client is built; unset is a Config error
.api_key_file(path)The file, when the client is built; surrounding whitespace is dropped

svir reads no environment variable and no .env file unless told to. The key is sent as Authorization: Bearer ... and never appears in Debug, Display, errors, or events. A local server without authentication needs no key at all.

Extra headers​

A gateway or a hosted endpoint may ask for a header of its own: attribution, an organization or a project, or a key under a name other than Authorization. header(name, value) adds one to every request the client sends, the model listing included.

use svir::prelude::*;

fn client() -> Result<Client, Error> {
Client::openai("https://gateway.example.com/v1")
.api_key_env("GATEWAY_KEY")
.header("x-title", "My App")
.build()
}
  • Names are not case-sensitive and are sent lowercase. Setting a name again replaces the earlier value.
  • Every value is treated as a credential: it never appears in Debug, in errors, or in events, and the built-in backend sends it marked sensitive. An error names the header, never its value.
  • svir reads no environment variable for a header. For a value that lives there, read it with std::env::var(..) and pass it.

build() fails with ErrorKind::Config for:

The headerWhy
A name that is not a header name ("x title", "")It cannot be sent
A value with a line break, another control character, or text that is not ASCIIA line break would end the header and start another
authorization, content-type, content-length, accept, host, transfer-encoding, connectionsvir writes these itself, or they frame the request. The Bearer key goes through api_key

The headers are the client's, the same on every request. There are no headers per request: a layer works above HTTP and cannot add one.

Timeouts​

WhatHowDefault
Connectingconnect_timeout(d)10 s
Silence from the server, before the response and during itidle_timeout(d)5 min
Time to the first piece of the answerTimeout::first_token(d) layerNone
The whole answerTimeout::total(d) layerNone

There is no total timeout by default, on purpose: a long generation is not a failure. The idle timeout is long because a local model reads the whole prompt before its first token. Shorten it for a hosted server, not for a local one.

All of them fail with ErrorKind::Timeout.

Listing models​

use svir::prelude::*;

async fn models(client: &Client) -> Result<Vec<String>, Error> {
let listed = client.list_models().await?;

Ok(listed.into_iter().map(|model| model.id).collect())
}

The listing is what the server reports, unfiltered: embedding and speech models are in it next to chat models. Model has id, the ID to put in a request, and an optional display name.

tip

Some local servers answer with whatever model is loaded when they do not know the ID in the request. If every model gives the same answer, list the models and check the ID.

Compatibility handling​

Some servers reject optional fields, reasoning_effort and stream_options, with a 400 or 422. The client then sends the request once more without them, and remembers: later requests of this client and its clones leave them out from the start. Nothing was generated by the rejected request, so this is not a retry in the sense of the Retry layer, and it needs no configuration.

What to know about it:

  • On such a server Completion::usage is None (usage is asked for through stream_options), and the reasoning effort is not applied.
  • Only a rejection whose body does not explain it gets the second attempt. A context overflow or a prompt the content filter blocked is reported as it is: sent again, a blocked prompt would be billed again.
  • If the second attempt fails too, the original error is reported.
  • A tool choice and a response format are not optional: the answer has to meet them. The second attempt keeps them, a request that carries them and no optional field is not sent twice, and a server that does not take them fails the request with Unsupported and its own words.