Skip to main content

Reading an answer

Every answer is streamed on the wire. You choose whether to read it as it arrives or wait for the end.

CallReturnsUse when
client.complete(&request).await?CompletionOnly the finished answer matters
client.stream(&request).await?EventStreamText is shown as it arrives
stream.completion().await?CompletionA stream was opened, and the rest of it is not shown

complete is stream followed by completion. stream resolves when the response starts, which for a local model can be long after the call: the model reads the whole prompt first.

complete and stream take the request by reference or by value.

Events​

EventStream::next() yields Option<Result<Event, Error>>. It is a method of the stream itself, so no StreamExt import is needed.

EventCarriesMeaning
Event::Text(String)A piece of the answerShow it
Event::Reasoning(Reasoning)source, textA piece of the model's reasoning; show it apart from the answer
Event::ToolCallDelta(ToolCallDelta)index, id, name, argumentsA piece of a tool call, for display only
Event::Completed(Completion)The whole answerAlways the last item

After Completed, or after an Err, next() returns None. Event is #[non_exhaustive]: end every match with _ => {}.

Show the deltas, keep the completion​

The completion already holds the whole answer, with its tool calls and usage. Use the deltas for display and keep what Completed carries; never rebuild the answer from the pieces.

use std::io::Write;

use svir::prelude::*;

async fn show(client: &Client, request: &Request) -> Result<Completion, Error> {
let mut stream = client.stream(request).await?;

while let Some(event) = stream.next().await {
match event? {
Event::Text(piece) => {
print!("{piece}");
let _ = std::io::stdout().flush();
}
Event::Reasoning(piece) => eprint!("{}", piece.text),
Event::Completed(done) => return Ok(done),
_ => {}
}
}

// Unreachable in practice: a stream ends with `Completed` or with an error.
Err(Error::new(ErrorKind::TruncatedStream))
}

EventStream is Send + Unpin + 'static and owns everything it needs, so it can be moved into a task or kept in a struct. It also implements futures_core::Stream, for code that wants combinators.

The completion​

FieldTypeHolds
finishFinishReasonStop, ToolCalls, Length, ContentFilter, or Refusal
textStringThe answer, exactly as sent
reasoningVec<Reasoning>Reasoning, one entry per source
callsVec<ToolCall>Complete tool calls, in order
usageOption<Usage>Token counts, when the server reported them
timingOption<Timing>When the first and last visible tokens arrived

Act on finish:

use svir::prelude::*;

fn describe(done: &Completion) -> &'static str {
match done.finish {
FinishReason::Stop => "the answer is complete",
// Run the tools in `done.calls` and ask again; see Tools.
FinishReason::ToolCalls => "the model is waiting for tool results",
// The output limit cut it off. With a reasoning model the text can be
// empty: the reasoning used the budget. Raise `max_tokens`.
FinishReason::Length => "the answer was cut off",
// The server's content filter stopped it, or flagged it after it was
// streamed. `done.text` holds what was sent, which may be what was
// flagged: withdraw what the user was shown.
FinishReason::ContentFilter => "the answer was filtered",
// The model would not answer, and `done.text` says why: show it as
// the answer. It does not have a format the request asked for.
FinishReason::Refusal => "the model refused",
_ => "a finish reason this code does not know yet",
}
}

A ContentFilter finish can come after the whole answer. Azure OpenAI's asynchronous content filter streams the answer before vetting it and reports a block afterwards, even after the model's own stop; svir makes that the finish. Text before a block may hold what was blocked in Azure's default mode too. So code that shows the deltas as they arrive takes the text down on this finish, rather than leaving it up with a note.

A Refusal finish means the model would not answer, and the text is its refusal. Chat Completions sends a refusal in a field of its own, in place of the answer; svir streams it as Event::Text like any answer, so a chat shows it with no code of its own. OpenAI's models refuse above all when asked for an answer in a format they will not give; see Structured output. A refusal the content filter stopped is ContentFilter.

An answer asked for as JSON is read into a type with done.parse::<T>(), which parses only an answer that finished with Stop; see Structured output.

The text is what the server sent. A server that separates reasoning often starts the answer with blank lines: trim for display, store as is.

Reasoning​

Servers carry reasoning in three ways, and svir reads all of them into Reasoning { source, text }:

ReasoningSourceWhere it was
ReasoningContentThe reasoning_content field
ReasoningThe reasoning field
Think<think>...</think> inside the answer text

Inline <think> tags are split out of the text by default, so reasoning is not shown as the answer even when the server has no reasoning parser. A tag cut in half by a chunk boundary is held back until the next chunk decides it. .think(Think::Keep) on the client builder leaves the tags in the text.

Completion::reasoning has one entry per source, in order of first appearance, with the pieces joined. Reasoning goes to the user as reasoning, never as the answer.

Usage and speed​

use svir::prelude::*;

fn report(done: &Completion) {
if let Some(usage) = done.usage {
println!("{} tokens in, {} out", usage.input, usage.output);

if let Some(reasoning) = usage.reasoning {
println!("{reasoning} of them reasoning");
}
}
if let Some(rate) = done.tokens_per_second() {
println!("{rate:.1} tokens per second");
}
}
  • Usage is asked for by default. A server that does not report it leaves usage as None; the answer is still complete. Never unwrap it.
  • usage.total and usage.reasoning are present only when the server sent them. svir reports what the server said; estimates are yours.
  • tokens_per_second() runs from the first visible token to the last, so the wait before the first token does not drag it down. It is None for a single token or a window under 50 ms: there is no honest rate then.
  • timing.first_token is measured from the start of the response, not from the call. Time the call yourself for time to first token as a user feels it.

Cancelling and deadlines​

Dropping the stream cancels the request and closes the connection. There is nothing to call. A cancellation token, a client that disconnected, a select! that moved on: all of them cancel by dropping.

use std::time::Duration;

use svir::prelude::*;

async fn bounded(client: &Client, request: &Request) -> Result<Completion, Error> {
let answer = tokio::time::timeout(Duration::from_secs(60), client.complete(request));

match answer.await {
Ok(result) => result,
// The future was dropped, and the request with it.
Err(_) => Err(Error::new(ErrorKind::Timeout).with_detail("no answer in 60 seconds")),
}
}

Conversely, a stream dropped early generated tokens nobody read.

For deadlines on every call of a client (time to first token, silence, the whole answer), use the Timeout layer. The client has one deadline built in: a server that sends nothing for 5 minutes fails the call.

When the stream fails​

An error is the last item of the stream. What arrived before it is a partial answer: fine to have shown, wrong to treat as complete, and its tool calls must not be run. A stream cut short never produces a completion at all.

KindWhat happened
TruncatedStreamThe connection closed before the answer finished
TimeoutThe server went silent, or a deadline passed
ServerThe server reported a failure inside the stream
ContextOverflowThe server said, inside the stream, that the request does not fit
Protocol, UnsupportedThe stream is malformed, or uses something svir does not read
ResponseLimitA limit was reached

svir retries none of this, and the Retry layer leaves a started answer alone: the user has already seen part of it. Sending the request again is the application's call. See Errors.