Skip to main content

Graceful Shutdown

A neva server stops on an OS signal — SIGINT / SIGTERM and the Windows equivalents — with no configuration at all. That is the right default for a process whose whole job is to be an MCP server, and no use for the two cases that are not that: a test that has to observe an orderly shutdown, and neva embedded in a larger service that owns its own lifecycle.

For those, App::with_shutdown(), with_shutdown_signal(..) and with_shutdown_drain(..) give you the stop explicitly — aborting the server task instead would skip every graceful path by construction.

Stopping without a signal​

with_shutdown() hands back a ShutdownHandle alongside the app:

use neva::prelude::*;

#[tokio::main]
async fn main() {
let (app, shutdown) = App::new()
.with_options(|opt| opt.with_default_http())
.with_shutdown();

let server = tokio::spawn(app.run());

// ... later, from anywhere:
shutdown.shutdown();

server.await.expect("the server task panicked");
}

The handle composes with the signal handler rather than replacing it — whichever fires first wins — so a server built this way still stops on Ctrl+C.

Shutdown is requested by the handle, not completed by it: shutdown() returns as soon as the request is recorded. Await run() to know the server actually finished.

Sharing a signal you already own​

Embedded in a service with its own lifecycle, take the signal from that service instead of handing one out:

use neva::prelude::*;
use neva::app::ShutdownHandle;

let shutdown = ShutdownHandle::new();
let app = App::new().with_shutdown_signal(shutdown.clone());

let server = tokio::spawn(app.run());
shutdown.shutdown();
server.await.expect("the server task panicked");
MethodDescription
ShutdownHandle::new()A handle backed by a fresh signal
ShutdownHandle::from_token(token)Wraps an existing tokio_util::sync::CancellationToken, so the server stops on a signal another subsystem already owns
handle.shutdown()Requests shutdown. Idempotent
handle.is_shutdown_requested()Whether shutdown was requested — not whether it finished

Clones share one signal: any clone calling shutdown() stops the server the handle came from.

What shutdown actually does​

Under MCP 2026-07-28 a server ending a subscription on its own initiative SHOULD send the empty result first, so a client can tell an orderly end from a dropped connection. Delivering it means shutdown is two-phase:

  1. The signal ends the subscriptions, and the server waits until the registry is empty and no message is still inside the middleware pipeline — together that means every result has reached the outbound channel.
  2. Only then is the transport torn down, and the writers drain what is queued before they exit.

with_shutdown_drain(..) caps the whole teardown:

use std::time::Duration;

App::new()
.with_shutdown_drain(Duration::from_secs(5))
.with_options(|opt| opt.with_default_http())
.run()
.await;
Default2 seconds
It is a ceiling, not a delayThe wait ends the moment the last result is queued, and is skipped outright when no subscription is open — a server that never uses them shuts down exactly as fast as it did before
Duration::ZEROOpts out, restoring an abrupt close

Raise it for a server whose subscriptions have deep buffers to flush.

The two halves share one budget

The deadline is stamped when the shutdown request arrives. Waiting for the subscriptions to answer spends part of it; the writers get the remainder. Whatever is still writing when it runs out is stopped rather than left on a runtime that may outlive the server.

Under run_blocking​

run_blocking() builds a runtime, runs the server on it, and drops the runtime the moment run returns — so anything still draining in a detached task would be aborted mid-write. run waits for the transport writers before returning, which is what makes the drain mean the same thing under both runners.

The HTTP engine has to stop too​

An HttpEngine is handed a CancellationToken and is expected to bring its listener down when the token fires. That contract is what the drain rests on: run waits for the engine's own run to return, so an engine that takes the token and never acts on it spends the whole shutdown budget on every stop.

Learn By Example​