Qwen3.8-27B for coding agents.
Flat-rate from $19/month.
Every request runs Qwen3.8-27B with FP8 weights and FP8 KV caching. Get predictable monthly billing, up to 256K context, and no routine prompt or response retention while keeping the chat-completions stack your tools already use.
- 78.8Btokens served
- 2.1Mrequests handled
Works with Pi · OpenClaw · Hermes · OpenCode · Cline · OpenAI SDK
Paid plans are for normal interactive use. Fair use, bounded queues, and shared-capacity availability apply.
Test for free, then choose predictable monthly access.
No token credits, no per-token overages, and no surprise usage invoice.
Free
Prove Yolo-Auto works in your stack before you pay.
- Fair use applies
- Designed for 1 coding agent
- Qwen3.8-27B
- 128K context for real prompts and docs
- No card required, free forever
Free forever. No card required and no trial period.
Builder
Best valueFor coding agents and daily development.
- No per-token billing
- Designed for roughly 3-4 coding agents
- 128K context for repositories and long conversations
- Best-effort capacity during busy periods
- Fair use applies
- Cancel anytime
Cancel anytime. No lock-in, no per-token overage.
Pro
For heavier interactive workloads that need more agent capacity and headroom.
- Everything in Builder
- Designed for roughly 5-6 coding agents
- 256K context for larger repositories
- More workload headroom
- Fair use applies
- Cancel anytime
Cancel anytime. No lock-in, no per-token overage.
Change three settings, not your stack.
Set the base URL, add a Yolo-Auto API key, and choose qwen3.8-27b. Keep the familiar request shape.
Focused model capacity for everyday agent work, without storing your prompts.
Works with clients that support a custom base URL and OpenAI-compatible chat completions. Run coding sessions, retries, tool calls, documents, and long-context work without changing request formats. We process request content to generate a response, not to build a conversation history or train models.
Pi, OpenClaw, Hermes, OpenCode, Aider, Cline, Roo Code, Continue, and other configurable clients.
OpenAI SDK clients, LangChain, LlamaIndex, curl scripts, and custom chat-completions integrations.
Use focused Qwen capacity for daily work and reserve frontier-model spend for tasks that genuinely need it.
Prompt and response bodies are not routinely retained as a browsable conversation archive.
Your requests and model responses are not fed into model-training pipelines.
We keep operational data such as model, token counts, status, and latency to run the service.
Narrow safety, security, abuse, and legal exceptions apply. Read the Privacy Policy.
Test the API before you pay.
Start with 15 requests per day. Upgrade when Yolo-Auto fits your stack and workload.
Start free →