Qwen3.8-27B API · OpenAI-compatible

Qwen3.8-27B for coding agents.
Flat-rate from $19/month.

Every request runs Qwen3.8-27B with FP8 weights and FP8 KV caching. Get predictable monthly billing, up to 256K context, and no routine prompt or response retention while keeping the chat-completions stack your tools already use.

  • 78.8Btokens served
  • 2.1Mrequests handled
Start free →15 requests/day · no card required Compare plans

Works with Pi · OpenClaw · Hermes · OpenCode · Cline · OpenAI SDK

Paid plans are for normal interactive use. Fair use, bounded queues, and shared-capacity availability apply.

Simple pricing

Test for free, then choose predictable monthly access.

No token credits, no per-token overages, and no surprise usage invoice.

Free

$0

Prove Yolo-Auto works in your stack before you pay.

  • Fair use applies
  • Designed for 1 coding agent
  • Qwen3.8-27B
  • 128K context for real prompts and docs
  • No card required, free forever

Free forever. No card required and no trial period.

Builder

Best value
$19/mo

For coding agents and daily development.

  • No per-token billing
  • Designed for roughly 3-4 coding agents
  • 128K context for repositories and long conversations
  • Best-effort capacity during busy periods
  • Fair use applies
  • Cancel anytime

Cancel anytime. No lock-in, no per-token overage.

Pro

$39/mo

For heavier interactive workloads that need more agent capacity and headroom.

  • Everything in Builder
  • Designed for roughly 5-6 coding agents
  • 256K context for larger repositories
  • More workload headroom
  • Fair use applies
  • Cancel anytime

Cancel anytime. No lock-in, no per-token overage.

Compare exact limits and estimated costs →
One endpoint

Change three settings, not your stack.

Set the base URL, add a Yolo-Auto API key, and choose qwen3.8-27b. Keep the familiar request shape.

POST /v1/chat/completions
Fits your workflow. Private by default.

Focused model capacity for everyday agent work, without storing your prompts.

Works with clients that support a custom base URL and OpenAI-compatible chat completions. Run coding sessions, retries, tool calls, documents, and long-context work without changing request formats. We process request content to generate a response, not to build a conversation history or train models.

Works withCoding-agent clients

Pi, OpenClaw, Hermes, OpenCode, Aider, Cline, Roo Code, Continue, and other configurable clients.

Works withSDKs and frameworks

OpenAI SDK clients, LangChain, LlamaIndex, curl scripts, and custom chat-completions integrations.

Built forPredictable cost

Use focused Qwen capacity for daily work and reserve frontier-model spend for tasks that genuinely need it.

PrivacyNo stored prompt history

Prompt and response bodies are not routinely retained as a browsable conversation archive.

PrivacyNo training on your requests

Your requests and model responses are not fed into model-training pipelines.

OperationsMetadata, not content

We keep operational data such as model, token counts, status, and latency to run the service.

Narrow safety, security, abuse, and legal exceptions apply. Read the Privacy Policy.

Test the API before you pay.

Start with 15 requests per day. Upgrade when Yolo-Auto fits your stack and workload.

Start free →