Skip to content

Models and context ​

ZettCode talks to any OpenAI-compatible endpoint: a hosted API, a gateway in front of several providers, or a model running on your own machine. What it needs to know is in Configuration.

Switching ​

/model opens a picker over the conversation. Enter selects for later requests — the one already running is not disturbed — and the header shows the new name immediately. /model DeepSeek Pro also works, matching either the display name or the model id, for when you know exactly what you want.

Each model keeps its own token, endpoint, and window, so switching between a local server and a hosted one is one keystroke rather than an edit and a restart.

For example, configure DeepSeek Pro and DeepSeek Flash with the two-model example, then try:

text
/model DeepSeek Pro
Explain the trade-offs in this module before making changes.
/model DeepSeek Flash
Read this screenshot and describe the error message.

The screenshot request needs an attached image and a model configured with multimodal = true. Choosing a model does not attach an image for you. The conversation remains the same when you switch; use /new for a clean start.

Reasoning effort ​

/effort picks how much the model should think before answering:

LevelWhat it is for
offRequest no deliberate reasoning, subject to the endpoint's support.
minimalA token of thought, fastest.
lowQuick answers for small edits.
mediumThe default balance.
highMore thought on hard problems.
xhighA long think on hard problems.
maxThe highest budget the provider accepts.
ultraHighest requested level; availability depends on the endpoint.

The choice applies to later requests for the current application run, and the header shows which one is selected. It is not a saved setting in config.toml. For example:

text
/effort low
Explain what this regular expression matches.
/effort high
Review this locking code for deadlocks and races; do not edit it yet.

The endpoint decides the actual reasoning budget. Not all endpoints support all levels, and a gateway may reject an unsupported value rather than map it. DeepSeek documents its current mapping in Thinking Mode. In this client, Chat Completions with /effort off omits reasoning_effort; it does not send DeepSeek's separate thinking toggle, so it is not a guarantee that a provider with thinking enabled by default will disable it.

The context window ​

Every model has a limit. ZettCode tracks how full it is — ctx 34.0% in the status line, and the full breakdown under /context — and compacts before it overflows, using compact_percent from the model's configuration.

Context  103,241 / 128,000 tokens
  Source              Share    Count
  System prompt        1.9%      (1)
  Environment notes    4.8%      (2)
  Tool schemas        10.9%      (9)
  User messages        0.1%      (3)
  Assistant messages  22.2%     (12)
  Tool output         60.1%     (17)
~4 chars/token · esc back

The title compares the total against the model's window; the percentages are shares of the context in use, so they add up to 100%. If a row is surprisingly large, that is the thing to look at: a tool schema you never use, or a long file that came back from a read. The footer names the counter that produced the numbers — tiktoken when its encoding could be loaded, a four-characters-per-token estimate otherwise.

Do not confuse the two percentages:

DisplayCalculationExample
Status ctxContext tokens ÷ configured window100,000 of 1,000,000 tokens is ctx 10.0%.
A /context rowThat source's tokens ÷ all current context tokens60,000 tool-output tokens out of 100,000 total is 60.0%.

The category shares add to 100% before display rounding. Count is the number of messages, tool definitions, or instruction items in that category, not the number of conversation turns. A resumed session can show a breakdown too. Before tools/environment notes have been assembled in the current process, the footer can say notes and tools pending; treat that snapshot as partial.

Compaction ​

When the newest request would cross the trigger, the conversation is summarized before it is sent: the older part becomes one summary message, and the newest quarter of the trigger stays verbatim, so the model keeps the immediate context and loses only what was already said. /compact does it immediately.

A summary is stored as its own kind of record, not as a delete. That is why the transcript still ends with Conversation checkpoint, why the models panel keeps working, and why a session compacted yesterday still resumes correctly. Asking by hand keeps only the last turn verbatim, where the automatic pass leaves a quarter of the trigger behind.

What compaction costs

Summarizing is one model call, so it takes a few seconds — the transcript shows Compacting while it runs. It is worth doing early rather than late: the alternative is an endpoint error at the moment you are in the middle of something. If a summary would not be smaller than the text it replaces, ZettCode says so and keeps the original.

Try this after several turns or a large file read:

text
/context
/compact
/context

Close each context panel with Esc before typing the next command. /compact starts the summary immediately and shows an animated Compacting row. Compare the second report with the first: older details are replaced by a summary, while the recent working context remains. This is not /clear: it changes what later model requests carry, not just what is visible on screen.

If the most recent checkpoint already covers the conversation, another manual compaction is refused with a notice. Continue the conversation before trying again. A summary may lose detail; ask the model to re-read a specific file when exact text matters.

Reading the numbers ​

↑18.4k ↓900 · 71.2% cached · 74 tok/s · ctx 12.3% — in order: tokens sent, tokens generated, the share of the input the provider served from its cache, generated tokens per second of model time, and how full the window was on the newest request. Input/output totals accumulate across the session; context occupancy is a snapshot, not an accumulated percentage.

FieldHow to read it
↑ / ↓Cumulative input and output tokens reported by the provider. Repeated context counts as input each time it is sent.
cachedCumulative cache-hit input tokens ÷ cumulative input tokens. 60,000 hits out of 100,000 input tokens is 60.0%.
tok/sOutput tokens divided by measured model time; not wall time spent on tools, approvals, or reading the answer.
ctxLatest context occupancy relative to the selected model's configured window.

An endpoint that does not report cache-hit usage cannot provide a meaningful hit-rate measurement here. These statistics are not a bill; check the provider's dashboard for charged usage.

The cache figure is worth understanding, because it is where time and money go. Providers cache a prefix of the request: if the beginning of a request is byte-for-byte identical to the previous one, that part is billed and served more cheaply. That is why the system prompt carries the date and not the time, project instructions are read in a fixed order, and /context is safe to open between turns — a stable prefix helps reuse cached input. There is no universal "good" percentage: model switches, compaction, changing tool definitions, and the provider's cache lifetime can all lower it. Newly generated output can only become cached input when a later request sends it back.

Released under the MIT license.