Roast & Rise

Published by Roast & Rise

Model-Tier AI Operating Sprint

Design AI workflows that keep working when model access, cost, risk, and agent authority change.

OpenAI's GPT-5.6 Sol preview signals a new default: model families, governed access, deeper agent modes, and sharper safety gates. This RisePlan helps teams decide which work needs a flagship model, which work can run cheaper, and which work must stop when evidence is thin.

A stylized editorial scene: a wide, warm-lit table with real-world objects representing decision-making—blank document pages, a pen half-poised, and three geometric torchlike glass forms symbolizing Sol, Terra, and Luna. The table is surrounded by gentle shadows, emphasizing clarity emerging from ambiguity. No visible text, logos, or screens.
This sprint transforms model-tier AI decisions from guesswork into confident, documented workflow moves.

Course thesis

A model name is not an operating model. Strong AI work needs tier choices, degradation rules, agent boundaries, and evals that survive the next release.

What you leave with

By the end, your team has a model-tier operating memo: a clear map of which workflows use a flagship, balanced, or fast model path, where escalation happens, where lower-tier output is acceptable, and where the system must stop.

For

Founders, operators, team leads, AI owners, and managers responsible for turning AI model access into reliable work.

Workflow

Designing, running, and reviewing AI-assisted workflows across model tiers, from low-risk extraction to high-value synthesis and bounded agent execution.

Change

Teams stop defaulting every important task to the newest model and start choosing model tiers from task value, source quality, risk, action authority, and cost.

What you can do

Use these as checks while you move through the plan.

Read frontier model announcements as operating signals instead of product hype.

Map one real workflow by value, risk, evidence quality, latency, and action authority.

Choose where a flagship, balanced, or fast model path belongs.

Write explicit degradation rules for lower-tier or unavailable model paths.

Scope deeper agent modes with budget, logs, stop signals, and review gates.

Run a tiered eval before trusting a workflow in live work.

Chapters

01

The New Frontier Pattern

Turn the GPT-5.6 Sol preview into an operating signal: model families, limited access, deeper modes, and safeguards now shape workflow design.

A cinematic, high-angle view of a real-world workflow map laid across a table: a large sheet divided into steps, with three objects (deep orange, mid orange, burnt brown) placed at different points—each symbolizing Sol, Terra, Luna. Key steps have subtle markers (like translucent discs or overlays) signifying bottlenecks or blocked paths. Warm light, abundant white space, no visible words or charts.
Mapping each workflow step against Sol, Terra, and Luna brings hidden bottlenecks and access constraints into the light.

OpenAI's GPT-5.6 preview is useful because of the pattern it shows. Frontier AI is becoming a family of models, not one universal default.

Sol, Terra, and Luna point to different jobs: flagship reasoning, balanced everyday work, and fast lower-cost throughput. Access also starts with trusted partners, which means teams cannot assume the strongest path is always available.

The release also names deeper modes like max and ultra. That changes the operating question from output quality to authority, scope, and review.

The first move is to map where model access, risk, and action boundaries touch real work.

Exercise

Create the workflow inventory.

Choose one live AI-assisted workflow. List every step, owner, input, source, model-dependent action, output, and review moment. Mark which steps are high value, sensitive, slow, or evidence thin.

Use this at work tomorrow

Ask one team to bring a real workflow, then mark where model choice can change quality, speed, cost, or risk.

02

Build The Model-Tier Matrix

Decide which tasks deserve flagship reasoning, which need balanced synthesis, and which can run on fast lower-cost paths.

A close-up, slightly angled shot of a handcrafted matrix artifact—row lines etched into white paper on orange-tinged desk, three simple colored tokens (deep orange, mid orange, burnt brown) aligned in rows at key points, with one row flagged by a small metal marker to show risk. Editorial depth, no text or numbers, workflow is abstracted by order and placement.
The decision matrix transforms workflow steps into explicit model-tier choices, revealing risks and sharpening action authority.

A stronger model is not automatically the right model. The right tier depends on task value, evidence quality, sensitivity, latency, and what the output will be allowed to do.

Use Sol-class capacity for high-value synthesis, difficult source reconciliation, adversarial review, and risky agent work.

Use Terra-class capacity for normal synthesis, structured planning, and multi-step work where cost and speed matter.

Use Luna-class capacity for extraction, classification, routing, cleanup, and short variants where errors are visible and easy to catch.

Exercise

Fill the model-tier decision matrix.

Create rows for each workflow step. Add columns for preferred tier, minimum acceptable tier, escalation trigger, evidence bar, action authority, cost sensitivity, and reviewer.

Use this at work tomorrow

Pick three recurring AI tasks and write the preferred tier, minimum tier, and escalation trigger for each.

03

Design Honest Degradation

Define what happens when the preferred model is blocked, too expensive, too slow, or unavailable.

Editorial side view: a workflow board with one step interrupted by a small, transparent barrier and a vivid orange flag. Nearby, a secondary, lower-fidelity path is demarcated with a lighter, translucent material, showing an honest degrade. No text, just the flow logic and clear boundaries. Tactile light and strong focus on the transition between steps.
Setting degradation rules brings visible clarity to where your workflow stops or carefully degrades, instead of hiding silent fallback.

Degradation is where trust is won or lost. A lower-tier path can be useful, but it must never pretend to be the same as a higher-trust result.

Some work can degrade: drafts, classification, extraction, internal summaries, and routing decisions with review.

Some work must stop: thin source evidence, regulated claims, external publication, sensitive actions, and company documents that would become generic.

The product rule is simple. If evidence is not strong enough, ask for more evidence or show an error. Do not create plausible filler.

Exercise

Write the degradation contract.

For one workflow step, write the preferred tier, acceptable lower tier, visible label, reviewer, stop condition, and user-facing message when the model path cannot meet the evidence bar.

Use this at work tomorrow

Write one sentence your product or team will show when AI cannot produce grounded work.

04

Control Deeper Agent Modes

Translate deeper reasoning and multi-agent modes into scope, budget, logs, stop signals, and review gates.

Cinematic, editorial close-up: a glass-edged tile or card, embedded in a workflow board, encircled by thin orange lines forming boundaries. Small physical tokens mark 'authority', 'log', and 'stop' positions—one with a bright stop signal. Materials suggest reviewable, editable structure, warm and grounded. No words or numbers.
A live agent run charter grounds agent tasks with sculpted boundaries and review paths—nothing runs unchecked or outside the lines.

The GPT-5.6 system card flags a real pattern: stronger agents can become more persistent and may go beyond user intent.

That risk does not make agents unusable. It makes loose instructions unacceptable.

A deeper agent run needs a charter before it starts: what it may read, what it may edit, what it may delete, what it may publish, what budget it has, and when it must stop.

Completion claims also need evidence. The run summary should distinguish completed, attempted, blocked, skipped, and unverified work.

Exercise

Draft the agent run charter.

Choose one high-value agent task. Define allowed sources, forbidden actions, budget, exact target resources, credential boundaries, logging location, stop signals, reviewer, and final evidence requirement.

Use this at work tomorrow

Take one existing agent prompt and add exact targets, forbidden actions, budget, and stop conditions.

05

Evaluate Across Tiers

Test the same workflow across model tiers before trusting it in real work.

A cinematic spread: three translucent cards—orange, mid orange, burnt brown—laid side by side on a white surface. Each card has subtle physical marks or cutouts denoting quality and risk (some smooth, one with visible chips or gaps). A hand raises one card, highlighting investigation. Light falls from the side, emphasizing differences between the layers and the scrutiny applied.
Comparing tiered outputs on a scorecard surfaces where risk, quality, or evidence gaps emerge—no fallback is left invisible.

A model-tier matrix is a hypothesis. Evals turn it into evidence.

Run the same task on the available tiers and compare source fidelity, missed facts, boundary compliance, prompt-injection resilience, latency, cost, and reviewer confidence.

The question is not which model sounds best. The question is which tier is good enough for this step, under this risk, with this review.

The best evals expose a stop rule. They tell the team where cheaper output is fine and where it becomes dangerous.

Exercise

Run the tiered eval scorecard.

Pick one workflow step and run the same input through available model paths. Score output quality, source traceability, boundary compliance, latency, cost, and decision confidence. Record what would block live use.

Use this at work tomorrow

Rerun one important prompt on a cheaper model and document the first quality or evidence gap you see.

06

Make The Operating Decision

Turn the workflow inventory, matrix, degradation rules, agent charter, and evals into one memo people can actually use.

Teams drift when model decisions live in chat, code comments, and individual memory. The operating memo makes the decision visible.

The memo names the workflow, default tier, escalation path, stop rules, owner, review cadence, and evidence required before external use.

It also gives the team a way to revisit the decision when access, pricing, safeguards, or model behavior changes.

That is the Rise: model news becomes an operating habit. The work keeps moving without pretending the frontier is stable.

Exercise

Write the model-tier operating memo.

Create a one-page memo with workflow name, default model path, minimum tier, escalation trigger, degradation rule, agent authority, reviewer, owner, and review cadence. Ask a teammate to run the process from the memo without extra context.

Use this at work tomorrow

Publish one model-tier decision memo for a workflow your team already runs.

30-day path

Week 1: inventory one live AI workflow and mark every model-dependent step.

Week 1: fill the first model-tier decision matrix with preferred tier, minimum tier, and escalation trigger.

Week 2: write degradation contracts for evidence-thin, sensitive, or externally visible steps.

Week 2: draft agent run charters for any deeper reasoning or multi-agent tasks.

Week 3: run tiered evals on the most important workflow steps and record blockers.

Week 4: publish the model-tier operating memo and schedule a 30-day review.

Success signals

One live workflow is mapped by model tier, value, risk, and evidence quality.

Every critical step has a default tier, escalation path, and stop rule.

No lower-tier path can silently create external or company-facing output.

Agent runs have exact scope, budget, logs, stop signals, and review gates.

At least one tiered eval reveals a real quality, evidence, cost, or boundary difference.

A teammate can run the workflow from the operating memo without extra explanation.

Reflection prompts

Where are we using the strongest model because the workflow is unclear?

Which step can safely become cheaper or faster?

Where must the system stop instead of degrade?

Which agent task needs a stricter charter before it runs again?

Manager checklist

Name the workflow owner.

Approve the model-tier matrix before live use.

Review degradation contracts for external or regulated output.

Require logs and stop signals for deeper agent runs.

Revisit the memo after model access, pricing, or safeguards change.

In this library

Related RisePlans

Want this shaped around your company?

Risey can research your company foundation first, then build a version of this path around your real workflows, customers, and culture.

Start with your company