ModelsSao10KL3-8B-Lunaris-v1-Turbo
providerSao10K /

L3-8B-Lunaris-v1-Turbo

14 DZD in 18 DZD out/ 1M tokens
Service tier pricing, in DZD per 1M tokens
TierInputOutputCached input
11.614.4—
Prices in DZD per 1M tokens

Lunaris is a merge of five Llama 3 models, made by one person, with the full recipe published — every source model, every density, every weight. Its creator describes the method candidly as black box magic, with values arrived at through years of personal experimentation rather than theory. It is built for roleplay and creative writing, and the recommended sampling reflects that: temperature 1.4, roughly ten times what most models in this catalogue suggest, held in check by a min-p floor that keeps the variance inside coherent choices. An 8,192-token window, and a model that reportedly holds its own against much larger ones at its task.

Publicfp8JSONStreamingRoleplayCreative
L3-8B-Lunaris-v1-Turbo
ArchitectureTransformer
Context Window8K

L3-8B-Lunaris-v1

A merge of five models, made by one person, with the whole recipe published.


Temperature 1.4

The number that defines how you use this model, and it is unlike anything else in this catalogue.

Model typeTypical recommendation
Instruction models0.15 – 0.7
Reasoning models0.6 – 1.0
Lunaris1.4

Roughly ten times the lowest recommendation here, and double the common one.

This is not an error. Temperature controls how much the model deviates from its most likely next token. Low temperature produces consistent, predictable text — which is what you want from an extraction pipeline and precisely what you do not want from a story.

A creative model sampled cold produces the obvious sentence every time. The character says the expected thing. The scene goes where you assumed. Nothing surprises anyone, including the person writing it.

Temperature 1.4 is the model being allowed to make unexpected choices.


And min_p 0.1 Is the Counterweight

The second half of the recommendation, and it is what makes the first half workable.

Temperature 1.4 alone would produce incoherence. At that setting, tokens with very low probability become reachable — and most low-probability tokens are low-probability because they are wrong.

min_p sets a floor relative to the top token. At 0.1, any token whose probability is below ten percent of the most likely token's is excluded entirely, regardless of temperature.

Read the two together:

min_p 0.1 decides which tokens are allowed — only ones the model considers genuinely plausible.

Temperature 1.4 decides how freely it chooses among them — flattening the distribution so the third-best option is a real possibility rather than a rounding error.

Controlled creativity, in two parameters. Variance inside a sensible set, rather than variance across everything.

Do not carry one without the other. Temperature 1.4 with a default min_p produces nonsense. min_p 0.1 with temperature 0.7 produces a duller model than the merge was tuned for.


Five Models, One Recipe

The merge is published in full — every source, every weight, every density.

Source modelDensityWeightWhat it contributes
Meta-Llama-3-8B-Instruct——Base
L3-8B-sunfall-v0.10.40.25Roleplay training
Jamet-8B-L3-MK10.50.3Roleplay and storytelling
badger-iota-llama-3-8b0.60.35General knowledge
Stheno-3.2-Beta0.70.4The predecessor

Method: ties. Precision: bfloat16. With int8_mask and rescale enabled, and normalize disabled.

Read the weights as a statement of intent. Stheno carries the most influence at 0.4 — this is explicitly an evolution of it. The general-knowledge merge sits second at 0.35, which is the deliberate counterweight: a roleplay model that has forgotten how the world works produces fluent scenes about nothing.

Publishing the recipe is the notable part. Most models describe what they are. This one shows how it was built, precisely enough to reproduce.


The Creator's Own Honesty

Worth quoting directly, because it is the most candid statement in this catalogue:

Merging seems to be black box magic though? In my personal experience merging multiple models from different datasets / data works better than combining them all in one. Values chosen are from long-running personal experimentation since Llama-2 Merging Era.

Three admissions in three sentences.

The method is not understood, and he says so rather than inventing a theory.

The approach is empirical — multiple models from different sources beat one combined training run, in his experience, without a claim about why.

The numbers come from years of trying things, not from a principle.

And a fourth statement elsewhere on the card is equally direct: "I personally think this is an improvement over Stheno v3.2." I personally think. Not "benchmarks show."

That honesty is worth more than a benchmark table would be here. Roleplay quality is not something a benchmark measures well, and a creator saying "this is my judgment from using it" is a more useful signal than a number that does not capture the thing you care about.


⚠️ 8,192 Tokens

The constraint that shapes every roleplay session on this model.

Llama 3 — not 3.1 — has an 8,192-token context. This merge inherits it.

On a creative writing model, that is the binding limitation. A roleplay session accumulates history: character definitions, scene setting, and every exchange since. At eight thousand tokens, that fills faster than it feels like it should.

What it holds. A character card, a scenario, and perhaps thirty to fifty exchanges depending on length. Not a novel. Not a long campaign without management.

Three habits that follow.

Keep the character definition tight. A two-thousand-token persona is a quarter of your window, charged on every turn.

Summarise rather than truncate. Dropping the oldest messages loses the beginning of the story silently; replacing them with a paragraph of summary keeps what happened while freeing the tokens.

Watch prompt_tokens, not message count. A few long exchanges consume more than many short ones.


The Llama 3 Instruct Template

Use the Llama 3 Instruct context template. Not ChatML, not Alpaca, not a custom format.

CODE
<|begin_of_text|><|start_header_id|>system<|end_header_id|>

{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>

{input}<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Through an API this is applied for you. It matters when self-hosting, and in front-ends where the template is a configuration choice — which is most roleplay front-ends.

A mismatched template on a merge is worse than on a base model. The merge was assembled from models that all speak Llama 3 Instruct; feeding it another format does not fail loudly, it produces a model that ignores your system prompt in ways that look like a quality problem.


Specifications

Model IDSao10K/L3-8B-Lunaris-v1-Turbo
Parameters8B
BaseMeta-Llama-3-8B-Instruct
Methodties merge of five models
Precisionbfloat16
Context8,192 tokens
TemplateLlama 3 Instruct
Recommended temperature1.4
Recommended min_p0.1
LicenceLlama 3 Community License
DeveloperSao10K

The licence is Meta's, inherited from the base model, with the conditions that carry — attribution and a monthly-active-user threshold among them.

Provider suffixes such as -Turbo are hosting conventions rather than the creator's naming, usually encoding a quantisation or a serving tier.


Capabilities

CapabilityValue
input_typestext
output_typestext
image_inputNot supported
context_window8192
reasoningNo separate reasoning trace
streamingSupported
tool_callingNot a design target
structured_outputNot a design target
requires_promptYes — text prompt required

This is a creative writing model. Tool calling and structured output are not what it was merged for, and expecting them is expecting something the recipe did not aim at.


Using Lunaris on DEVUP AI

Base URL: https://api.devupai.com/v1 · Model ID: Sao10K/L3-8B-Lunaris-v1-Turbo

Python

PYTHON
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEVUP_API_KEY"],
    base_url="https://api.devupai.com/v1",
)

response = client.chat.completions.create(
    model="Sao10K/L3-8B-Lunaris-v1-Turbo",
    messages=[
        {"role": "user", "content": "Hello world!"}
    ],
    max_tokens=1024,
)

print(response.choices[0].message.content)

Node.js

JAVASCRIPT
import DevupAI from "devupai";

const client = new DevupAI({
  apiKey: process.env.DEVUP_API_KEY,
});

async function main() {
  const response = await client.chat.completions.create({
    model: "Sao10K/L3-8B-Lunaris-v1-Turbo",
    messages: [{ role: "user", content: "Hello world!" }],
    max_tokens: 1024,
  });

  console.log(response.choices[0].message.content);
}

main();

cURL

BASH
curl -X POST "https://api.devupai.com/v1/chat/completions" \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Sao10K/L3-8B-Lunaris-v1-Turbo",
    "messages": [
      { "role": "user", "content": "Hello world!" }
    ],
    "max_tokens": 1024
  }'

At the Creator's Settings

The configuration this model was tuned for, rather than the defaults.

PYTHON
response = client.chat.completions.create(
    model="Sao10K/L3-8B-Lunaris-v1-Turbo",
    messages=[
        {"role": "system", "content": SCENE},
        *history,
        {"role": "user", "content": turn},
    ],
    temperature=1.4,
    max_tokens=512,
    extra_body={"min_p": 0.1},
)

Both parameters, together. Temperature 1.4 without the min-p floor produces incoherence; the floor without the temperature produces a duller model than the merge was assembled to be.

⚠️ min_p is not universally supported through an OpenAI-compatible interface. Confirm it reaches the model on your path — and if it does not, lower the temperature toward 1.0 rather than running 1.4 unguarded.

512 output tokens is a reasonable turn length for roleplay. Longer responses eat the 8,192-token window faster, and a model that writes four paragraphs when one would do is a model that fills your context with its own prose.


Writing the Scene

On a creative model, the system prompt is the scene rather than a configuration.

PYTHON
SCENE = """You are writing collaboratively with the user.

Setting: a rain-soaked port city, late autumn, some decades after a war nobody won.

Your character: Idir, harbourmaster. Fifty-one. Keeps meticulous records and an unregistered ledger
alongside them. Speaks plainly, rarely finishes a sentence about the war. Does not trust anyone who
arrives by land.

Rules:
- Write Idir only. Never write the user's character's words, thoughts, or actions.
- Two to four paragraphs. Leave the user something to respond to.
- Stay in the scene. No summaries, no meta-commentary, no asking what happens next."""

Four things that matter on a roleplay model.

Never write the other character. The most common failure, and the most disruptive — a model that narrates the user's actions has taken the story away from them.

Set a length. Without it, a creative model tuned hot will write until it runs out of budget.

Forbid meta-commentary. "What would you like to do next?" breaks the scene, and models do it when uncertain.

Give the character a contradiction. Meticulous records and an unregistered ledger. That is what makes a character behave interestingly rather than consistently — and a merge tuned for creativity will use it.


Managing an 8,192-Token Session

The mechanism a long session needs.

PYTHON
SUMMARY_TRIGGER = 6000   # leave room for the response
KEEP_RECENT = 8          # recent turns stay verbatim


def compact(history: list[dict], scene: str) -> list[dict]:
    """Summarise older turns when the conversation approaches the window."""
    estimated = sum(len(m["content"]) for m in history) // 3

    if estimated < SUMMARY_TRIGGER:
        return history

    old, recent = history[:-KEEP_RECENT], history[-KEEP_RECENT:]

    transcript = "\n\n".join(f"{m['role']}: {m['content']}" for m in old)

    summary = client.chat.completions.create(
        model="Sao10K/L3-8B-Lunaris-v1-Turbo",
        messages=[
            {
                "role": "system",
                "content": (
                    "Summarise this story so far in under 200 words. Keep: what happened, what each "
                    "character learned, what remains unresolved, and any promise or threat made. "
                    "Drop: description, atmosphere, dialogue that led nowhere."
                ),
            },
            {"role": "user", "content": transcript},
        ],
        temperature=0.4,
        max_tokens=400,
    ).choices[0].message.content

    return [{"role": "system", "content": f"Story so far:\n{summary}"}, *recent]

Note the temperature drop to 0.4 for the summary. Summarising is not creative writing — it wants the obvious sentence, which is exactly what low temperature produces.

The instruction about what to keep is the important part. A summary that preserves atmosphere and loses a promise made three scenes ago will produce a story that contradicts itself.

Trigger before you hit the wall. At 6,000 tokens there is still room to generate the summary; at 8,000 there is not.


Why It Reportedly Beats Larger Models

A claim that circulates about Lunaris — that it holds its own against models in the 15B to 70B range at roleplay — and it is worth understanding rather than accepting.

Roleplay quality is not the same axis as general capability. A 70-billion-parameter general model knows more, reasons better, and follows complex instructions more reliably. It may also write characters who sound like a helpful assistant playing a character.

A merge assembled specifically from roleplay-trained models inherits what those models learned about dialogue, pacing, and staying in a voice — at eight billion parameters.

So the comparison is real and narrow. Better at this. Not better generally.

And the honest caveat: that claim comes from community assessment rather than measurement, because roleplay quality resists measurement. Try it on your own scenes; that is the only evaluation that means anything here.


⚠️ Content and Moderation

Worth stating plainly.

This model is assembled from roleplay-focused fine-tunes, and models in that category are generally tuned to be less restrictive than instruction models from major labs. That is what makes them useful for fiction; it is also what makes them different to operate.

Three consequences for a public-facing deployment.

The guardrails are yours. There is no lab-tuned refusal behaviour to rely on, and none of the source models were built with one.

Filter both directions. Input filtering catches intent; output filtering catches what the model produced, which on a creative model sampled at temperature 1.4 is not always what the prompt suggested.

Write your own content policy into the system prompt, and treat it as a prompt rather than a guarantee — a model tuned for creative freedom will follow a scene's momentum.

A dedicated moderation model is the right tool for the filtering, and there are several in this catalogue built for exactly that.


Where It Fits

Collaborative fiction and roleplay, which is what it was merged for.

Creative writing assistance — character development, dialogue, scene drafting, working through a block.

Character-driven applications, where a distinct voice matters more than factual precision.

Fast, cheap creative generation at eight billion parameters, where a larger model's cost per turn would not survive a long session.

Not for factual work. Eight billion parameters, merged for creativity, sampled at temperature 1.4 — every one of those pushes away from accuracy.

Not for structured output or tool use. Not what the recipe aimed at.

Not for long sessions without compaction. 8,192 tokens fills quickly.

Not for public deployment without your own moderation layer.


Practical Notes

Use temperature 1.4 and min_p 0.1 together. Neither works properly without the other.

Confirm min_p reaches the model on your path; lower the temperature if it does not.

Use the Llama 3 Instruct template.

Never let the model write the user's character. Say so explicitly.

Set a response length, or it will use the budget you gave it.

Compact the history at 6,000 tokens rather than waiting for the wall.

Drop the temperature for summarisation — that is not creative work.

Build a moderation layer for any public path.

Evaluate on your own scenes. Benchmarks do not measure what this model is for.


Limitations

8,192-token context. Short by any current standard, and the binding constraint on a roleplay session.

Eight billion parameters. Strong at its narrow task and not a capable general model.

Temperature 1.4 is not a factual setting. This model is configured for variance, and variance is the opposite of accuracy.

Text only. No image, audio, or video input.

No tool calling or structured output as design targets.

A merge, and the creator says the method is not understood. That candour is a virtue and it is also a statement about predictability — merged models can behave unexpectedly in ways a trained model does not.

Quality claims are community assessment, not measurement. Roleplay quality resists benchmarking, which cuts both ways.

The Llama 3 Community License, not a permissive one.

Built on Llama 3, released mid-2024. The field has moved, and newer creative models exist with substantially longer context.

Content moderation is entirely your responsibility. The source models were not built with one, and a model tuned for creative freedom follows a scene wherever it goes.