Speaks OpenAI-compatible Chat Completions, as served by
- LM Studio
- llama.cpp
- vLLM
- mlx-lm
- Azure OpenAI
- OpenAI-compatible endpoints
Streaming, end to end
The request body streams from disk with an exact length; the answer streams back as events. The last event is the whole answer, and dropping the stream cancels the request.
Nothing lost silently
Tool-call IDs and reasoning survive a round trip. Something that cannot be represented is an explicit error, never a dropped field. Keys never reach errors, events or logs.
Tools without macros
A tool is a description and a handler. Arguments deserialize into your type, schemas can be derived from it, and the loop that feeds results back stays in your hands.
Layers
Middleware around every call: Retry, Timeout and Trace built in, a closure or a type of your own next to them. A client without layers pays nothing for them.
Strict and bounded
Unknown input is an error unless you ask for leniency. Bytes, events and tool calls always have limits, and reaching one is a typed outcome, not a hang.
Testable without a model
The transport sits behind one trait. Swap in a scripted backend and test the code that calls a model with no server, no network and no keys.
Stream an answer
Text arrives as events while the model writes it. The last event carries the whole answer: text, reasoning, tool calls, usage and timing. client.complete waits for it instead, and dropping the stream cancels the request.
use svir::prelude::*;
#[tokio::main]
async fn main() -> Result<(), Error> {
let client = Client::openai("http://127.0.0.1:1234").build()?;
let request = Request::new("qwen3-27b")
.system("Be precise.")
.user("Why do rivers meander?");
let mut stream = client.stream(&request).await?;
while let Some(event) = stream.next().await {
match event? {
Event::Text(piece) => print!("{piece}"),
Event::Completed(done) => println!("\n{:?}", done.usage),
_ => {}
}
}
Ok(())
}
Give it tools
A tool is a description and a handler. Arguments are deserialized into your type, the schema can be derived from it, and a failure goes back to the model as text it can act on. The loop that feeds results back is a few lines of your own, with a bound you choose.
Toolsuse schemars::JsonSchema;
use serde::Deserialize;
use svir::prelude::*;
#[derive(Deserialize, JsonSchema)]
struct City {
/// The city to look up.
city: String,
}
async fn weather(args: City) -> Result<String, String> {
match args.city.to_lowercase().as_str() {
"oslo" => Ok("4 C, light snow".to_owned()),
_ => Err(format!("no weather station in {}", args.city)),
}
}
async fn ask(client: &Client, question: &str) -> Result<String, Error> {
let tools = Tools::new().add("weather", "The weather in a city right now.", weather);
let mut request = Request::new("qwen3-27b").tools(&tools).user(question);
for _ in 0..8 {
let done = client.complete(&request).await?;
if done.calls.is_empty() {
return Ok(done.text);
}
let results = tools.call_all(&done.calls).await;
request = request.assistant(done).tool_results(results);
}
Err(Error::new(ErrorKind::Unsupported).with_detail("the model kept calling tools"))
}
Compose layers
Middleware around every call. Retry, Timeout and Trace are built in; a closure or a type of your own goes in the same chain. Nothing is retried once the answer has started.
use std::time::Duration;
use svir::layer::{Retry, Timeout, Trace};
use svir::prelude::*;
fn client() -> Result<Client, Error> {
Client::openai("https://models.example.com/v1")
.api_key_env("MODELS_API_KEY")
.layer(Retry::transient(3).backoff(Duration::from_millis(250)))
.layer(Timeout::first_token(Duration::from_secs(60)))
.layer(Trace)
.wrap(|request, next| async move {
eprintln!("asking {}", request.model);
next.run(request).await
})
.build()
}
Relay through a proxy
Pass the server's bytes on unchanged and read them on the way past. The push-based decoder does no I/O, so a chat backend can stream to a browser and still keep the finished answer.
Relaying through a proxyuse bytes::Bytes;
use svir::openai::chat::Decoder;
use svir::prelude::*;
/// Relays the answer through `forward`, and returns it once it is complete.
async fn relay(
client: &Client,
request: &Request,
mut forward: impl FnMut(Bytes),
) -> Result<Completion, Error> {
let mut raw = client.send(request).await?;
let mut decoder = Decoder::lenient();
let mut answer = None;
while let Some(bytes) = raw.next().await {
let bytes = bytes?;
forward(bytes.clone());
for event in decoder.push(&bytes) {
if let Event::Completed(done) = event? {
answer = Some(done);
}
}
}
decoder.finish()?;
answer.ok_or_else(|| Error::new(ErrorKind::TruncatedStream))
}