Skip to main content

svir

A small, composable Rust SDK for talking to large language models.

The wire protocol between your application and a model server, and nothing it does not need.

Speaks OpenAI-compatible Chat Completions, as served by

  • LM Studio
  • llama.cpp
  • vLLM
  • mlx-lm
  • Azure OpenAI
  • OpenAI-compatible endpoints

Streaming, end to end

The request body streams from disk with an exact length; the answer streams back as events. The last event is the whole answer, and dropping the stream cancels the request.

Nothing lost silently

Tool-call IDs and reasoning survive a round trip. Something that cannot be represented is an explicit error, never a dropped field. Keys never reach errors, events or logs.

Tools without macros

A tool is a description and a handler. Arguments deserialize into your type, schemas can be derived from it, and the loop that feeds results back stays in your hands.

Layers

Middleware around every call: Retry, Timeout and Trace built in, a closure or a type of your own next to them. A client without layers pays nothing for them.

Strict and bounded

Unknown input is an error unless you ask for leniency. Bytes, events and tool calls always have limits, and reaching one is a typed outcome, not a hang.

Testable without a model

The transport sits behind one trait. Swap in a scripted backend and test the code that calls a model with no server, no network and no keys.

01

Stream an answer

Text arrives as events while the model writes it. The last event carries the whole answer: text, reasoning, tool calls, usage and timing. client.complete waits for it instead, and dropping the stream cancels the request.

Reading an answer
use svir::prelude::*;

#[tokio::main]
async fn main() -> Result<(), Error> {
let client = Client::openai("http://127.0.0.1:1234").build()?;
let request = Request::new("qwen3-27b")
.system("Be precise.")
.user("Why do rivers meander?");

let mut stream = client.stream(&request).await?;
while let Some(event) = stream.next().await {
match event? {
Event::Text(piece) => print!("{piece}"),
Event::Completed(done) => println!("\n{:?}", done.usage),
_ => {}
}
}

Ok(())
}
02

Give it tools

A tool is a description and a handler. Arguments are deserialized into your type, the schema can be derived from it, and a failure goes back to the model as text it can act on. The loop that feeds results back is a few lines of your own, with a bound you choose.

Tools
use schemars::JsonSchema;
use serde::Deserialize;
use svir::prelude::*;

#[derive(Deserialize, JsonSchema)]
struct City {
/// The city to look up.
city: String,
}

async fn weather(args: City) -> Result<String, String> {
match args.city.to_lowercase().as_str() {
"oslo" => Ok("4 C, light snow".to_owned()),
_ => Err(format!("no weather station in {}", args.city)),
}
}

async fn ask(client: &Client, question: &str) -> Result<String, Error> {
let tools = Tools::new().add("weather", "The weather in a city right now.", weather);
let mut request = Request::new("qwen3-27b").tools(&tools).user(question);

for _ in 0..8 {
let done = client.complete(&request).await?;
if done.calls.is_empty() {
return Ok(done.text);
}

let results = tools.call_all(&done.calls).await;
request = request.assistant(done).tool_results(results);
}

Err(Error::new(ErrorKind::Unsupported).with_detail("the model kept calling tools"))
}
03

Compose layers

Middleware around every call. Retry, Timeout and Trace are built in; a closure or a type of your own goes in the same chain. Nothing is retried once the answer has started.

Layers
use std::time::Duration;

use svir::layer::{Retry, Timeout, Trace};
use svir::prelude::*;

fn client() -> Result<Client, Error> {
Client::openai("https://models.example.com/v1")
.api_key_env("MODELS_API_KEY")
.layer(Retry::transient(3).backoff(Duration::from_millis(250)))
.layer(Timeout::first_token(Duration::from_secs(60)))
.layer(Trace)
.wrap(|request, next| async move {
eprintln!("asking {}", request.model);
next.run(request).await
})
.build()
}
04

Relay through a proxy

Pass the server's bytes on unchanged and read them on the way past. The push-based decoder does no I/O, so a chat backend can stream to a browser and still keep the finished answer.

Relaying through a proxy
use bytes::Bytes;
use svir::openai::chat::Decoder;
use svir::prelude::*;

/// Relays the answer through `forward`, and returns it once it is complete.
async fn relay(
client: &Client,
request: &Request,
mut forward: impl FnMut(Bytes),
) -> Result<Completion, Error> {
let mut raw = client.send(request).await?;
let mut decoder = Decoder::lenient();
let mut answer = None;

while let Some(bytes) = raw.next().await {
let bytes = bytes?;
forward(bytes.clone());

for event in decoder.push(&bytes) {
if let Event::Completed(done) = event? {
answer = Some(done);
}
}
}
decoder.finish()?;

answer.ok_or_else(|| Error::new(ErrorKind::TruncatedStream))
}

The name

The Svir is the river that joins Lake Onega to Lake Ladoga; the Neva then carries that water to the Baltic. svir sits in the same family as volga (HTTP) and neva (MCP), and like its river it joins two bodies of water that already exist: your application and a model server.