Altworld Astrea: The New Frontier of Creative Writing

Try Astrea → · Get an API key · Read the docs →

Today, we are introducing Altworld Astrea — our first openly released model. Astrea is a compact 9B-parameter model built for one thing: writing that feels written. It sounds human, it keeps the story straight, and it beats bigger models at both. You say something, it answers back.

Altworldbench — overall writing quality. Elo from 794 blind matchups on our release pool — both rivals are bigger models built for creative writing:

Elo
Astrea 9B 1114
Rocinante-X 12B 1097
Wayfarer-2 12B 1045

Astrea is available today on chat.altworld.io, and the full model weights are right here under Apache-2.0, so anyone can download it and run it locally. An OpenAI-compatible API is open with simple prepaid pricing.

What makes Astrea different

It sounds like a person

Most AI prose gives itself away in the first paragraph: the stock phrases, the "it wasn't X — it was Y", the em-dash pileups, every reply opening the same way. We count those tells. Astrea scores 2.5 — roughly two and a half tells slipping into every thousand words. Every other model we tested leaves between 8 and 19 per thousand.

It understands the scene, not just the sentence

Judged blind, story against story, Astrea wins on the things that make writing actually work: characters whose motivations hold together, a detail planted early paying off later, the tone you asked for instead of generic fantasy voice. That depth is what puts it ahead of models three billion parameters bigger.

It stays grounded

If the lantern went out, it stays out. If your sword was taken, it doesn't reappear in your hand two scenes later. That's grounding — remembering what's true in the story — and creative models are notoriously bad at it. Astrea keeps 73 of 100 scenes straight. The strongest 12B prose model manages 44.

All of it at 9B

Every rival on our board is both bigger than Astrea and purpose-built for creative writing. Astrea isn't beating generalists — it's beating the specialists, with three billion fewer parameters. That also means it runs faster, serves cheaper, and fits on hardware the bigger models can't.

Altworldbench

We test every release on Altworldbench, our own creative-writing benchmark. Same prompts and settings for every model, every story judged blind.

Head-to-head wins — share of blind matchups Astrea wins:

wins
vs Rocinante-X 52%
vs Wayfarer-2 65%

AI-sounding writing — lower is better:

score
Astrea 9B 2.5
Rocinante-X 8.1
Wayfarer-2 8.3
Muse 8.8
Violet Lotus 12.8

Keeps the story straight — out of 100:

score
Astrea 9B 73
Violet Lotus 68
Muse 64
Wayfarer-2 61
Rocinante-X 44

Writes the length you ask for — % of the time:

%
Astrea 9B 93
Rocinante-X 88
Violet Lotus 83
Wayfarer-2 78
Muse 25

All rivals are open 12B creative-writing models.

See for yourself

Two samples from our test set — unedited, exactly as Astrea wrote them.

SCENE: The rescue party returns with empty stretchers and one boot. ACTION: I ask the search captain to tell the families myself.

The captain sets the boot on the table. It is a heavy boot, the kind the welders wear, with the steel cap dented in and the laces frayed. The stretchers are empty now, returned to the hangar. The rescue party came back in four, and one of them did not return.

The families are waiting in the corridor. They have not eaten in days. The ration scrip is worthless now that half the station is gone, and the spin-section has slowed to a crawl. The air smells of recycled carbon and old sweat. You do not speak of the lost. You do not speak of the dead.

The captain turns to you. He does not speak. The boot is on the table. He steps back.

You stand. You take the boot. You walk into the corridor. The families look at you. There is nothing to say. You set the boot on the floor and say, "Your children did not come back."

SCENE: Your partner swore your cargo of walrus ivory to two different buyers. ACTION: I go to the lesser buyer first and offer him the truth and a discount.

The lesser buyer squints at you through a greasy half-door in Uppsala's timber shed. The air smells of pine pitch and old fish.

He does not lie. The other party is already buying that ivory. Three measures from Helsingjar, not yours. You tell him the truth, offer him a discount of two dirhams. He laughs, spits. "The other man has my seal. He paid me in hacksilver and a vow over the sacred waters. You offer scraps, and I must burn them for luck."

He turns away. The floorboards are wet. Frost lies thin on Lake Mälaren beyond the sheds, and the ice does not hold for a barge. The cargo remains in the gloom: twenty tusks, the colour of a seal's eye, unsworn to you now.

Availability

  • Talk to Astrea — chat with it now at chat.altworld.io — no account needed for your first turns.
  • Build with the API — an OpenAI-compatible endpoint at api.altworld.io — create a key and load credits at altworld.io/developer, then start in the docs.
  • Play Astrea — Astrea powers the voices of Altworld, our life simulator and story sandbox.
  • Download the weights — this repository: Apache-2.0, BF16 safetensors.

Build with the API

The Astrea API is OpenAI-compatible: if your code already talks to an OpenAI-style endpoint, point it at our base URL, swap the model name, and it works — streaming included.

curl https://api.altworld.io/v1/chat/completions \
  -H "Authorization: Bearer $ASTREA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "astrea",
    "messages": [{"role": "user", "content": "Open a story in a lighthouse, first person, no cliches."}],
    "stream": true
  }'

Model details

Altworld Astrea (release name Astrea-R8-Chat-9B) was trained on top of the Qwen3.5-9B base model using multi-round LoRA supervised fine-tuning, alternating conversational and narrative rounds and checking every round against Altworldbench. The released checkpoint is the merged BF16 model.

The chat template works bare: no system prompt, no flags, it just answers — with no visible reasoning block. Explicit thinking is available per-request with enable_thinking=True.

Recommended sampling:

"temperature": 0.8,
"min_p": 0.025,
"repetition_penalty": 1.08

Use temperature 0.2 for factual chat.

Limitations

  • Like any compact model, Astrea can miss nuance or state things confidently and wrongly. Verify anything high-stakes independently.
  • English dominates its training and evaluation. Other languages will work, but with weaker guarantees.
  • This release is text-first: the architecture inherits vision support, but visual behavior was not part of the release evaluation.
  • Astrea is not certified for medical, legal, financial, or other high-stakes use.

We plan to keep improving and openly releasing Astrea models — tell us what to fix. The fastest way to reach us is the Altworld Discord, or open a Discussion on this repository.


The fine print

Evaluation protocol. The pooled Elo is 1113.7 with a 95% bootstrap interval of 1089.1–1140.9 on a frozen seven-model pool (794 comparisons for Astrea's entry). The interval overlaps Astrea R8 and Rocinante-X-12B-v1, so the supported claim is #1 on point estimate and statistical parity with the top cluster — not a statistically separated win. Pairwise judgments were made blind by GLM-5p1 and DeepSeek-V4-Flash, with Kimi-K2.6 adjudicating disagreements; grounding, slop, length/POV/tense, and bare-chat checks combine deterministic rules with LLM judging. Separately from the narrative pool: 92.27% pass on 207 held-out bare-chat prompts, 0.00% accidental fiction, 20/20 on unseen identity paraphrases. One constructed multi-turn memory probe answered indirectly at temperature 0.2 while fine at 0.8 — test your own dialogue distributions.

Running it locally. Validated in production on vLLM 0.25 (A40, and 2×RTX 3090 with --tensor-parallel-size 2 — set NCCL_P2P_DISABLE=1 on consumer dual-GPU boards):

vllm serve Altworld/Astrea-R8-Chat-9B \
  --dtype bfloat16 --max-model-len 32768 --gpu-memory-utilization 0.90

The base architecture natively supports 262k context; the hybrid linear-attention layers keep the KV cache small, so serve what your VRAM affords. With Transformers, load via AutoProcessor + AutoModelForMultimodalLM and apply the bundled chat template (enable_thinking defaults off). An MLX-LM smoke run behaved worse than vLLM at temperature 0.2 despite numerically matched logits — validate your exact engine. GGUF quants are planned; llama.cpp supports the Qwen3.5 family.

Lineage and license. Base model Qwen/Qwen3.5-9B, Apache-2.0; this derivative is released under the same license — see NOTICE. Fine-tuning preference data used gold responses from MIT-licensed open-weight models (GLM-5.2, DeepSeek-V4-Flash) with on-policy rejected sides; no closed-model outputs anywhere. Early-stage data ancestry credits: Athanorlite-DPO (CC-BY-4.0), jondurbin/gutenberg-dpo-v0.1 (CC-BY-4.0), nbeerbower/gutenberg2-dpo and gutenberg-moderne-dpo (Apache-2.0), Pageshift-Entertainment/LongPage, allenai/tulu-3-sft-mixture (ODC-BY, license-filtered), nothingiisreal/Reddit-Dirty-And-WritingPrompts (Apache-2.0, SFW file only). Thanks to the Qwen team for the base model.

Downloads last month
912
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Altworld/Astrea-R8-Chat-9B

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(528)
this model
Quantizations
4 models