
Claude Fable 5.5 vs Fable 5.1: How to Test an Upgrade
What is known about the upgrade path?
| Migration question | What you can do now | What must wait for evidence |
|---|---|---|
| Can the model ID be changed directly? | Find where the current route is configured | Confirm the exact candidate ID and supported request contract |
| Will tools and structured outputs behave the same? | Preserve schemas, fixtures and acceptance checks | Replay against a verified candidate route |
| Will existing conversations continue correctly? | Save representative sanitized histories | Test history acceptance and continuation behavior |
| Will caching and costs carry over? | Record current usage and actual charges | Verify candidate rules and billing; assume no cache portability |
| Is a full replacement necessary? | Identify workloads that actually need improvement | Compare outcomes and decide whether to migrate at all |
The rest of this guide is a proposed migration procedure. It does not imply that Fable 5.5 supports any particular feature or that a release is scheduled.
Start with Fable 5.1's actual integration constraints
tool_choice modes any and tool as errors. It also documents thinking-block restrictions when moving to older models or changing earlier conversation content. These are existing 5.1 rules, not newly discovered 5.5 changes; verify their applicability to the exact gateway route separately.| Existing dependency | What to preserve from the current app | What a future migration must answer |
|---|---|---|
| Tool selection | Tool schemas, required business action and how the app verifies it actually occurred | Does the candidate support the chosen control, and can it complete the action without relying on an unsupported forced choice? |
| Thinking-bearing history | Original ordered messages and opaque blocks as received, plus the history-transform version | Which blocks remain valid on the candidate and on the fallback model? |
| Client-side compaction | Before/after transcripts, which turns were summarized, which later blocks were retained | Does the transformed history remain accepted and preserve the user's decisions? |
| Streaming and parsing | Raw event samples, tool-call assembly, termination/error branches and parser version | Can the existing parser reconstruct the required output and distinguish completion from interruption? |
| Usage and cache | Non-overlapping usage categories, actual charges, cold and warm runs | Does the candidate preserve the expected cache behavior, and what is the new total session cost? |
This creates an important diagnostic distinction: a 5.1 request rejected after your app rewrites earlier turns is already a baseline integration problem. It should not be counted as a regression introduced by an untested successor. Run each fixture on the current working route first; document any failure before adding the candidate.
Business state and model reasoning state also need separate recovery paths. A ticket ID and a user-approved owner belong in your application's durable state. Thinking blocks are model-specific transcript artifacts; keeping them does not guarantee that another model can consume them. A fallback can retain the approved business facts yet require a different, documented history representation.
Avoid changing the model, system prompt, tools and compaction scheme in one experiment. Save prompt templates, tool definitions, representative tool responses and parser versions. Mark which outputs software consumes and which a person reviews. First replay in a sandbox or with recorded tool responses: sending a message, creating a record or charging an account twice is not an acceptable side effect of a comparison.
Build a replay set that can detect regressions
Include successful Fable 5.1 sessions as well as unresolved failures. A candidate that repairs one hard task but disrupts common workflows may not be an upgrade for your application. Define the expected outcome before examining the candidate output.
| Replay case | Check | Example failure condition |
|---|---|---|
| Fresh request | Required instructions and output fields are respected | A required field disappears or a constraint is ignored |
| Long existing history | Important state and user decisions remain intact | The continuation contradicts an earlier accepted decision |
| Tool call and tool result | Arguments, sequencing and final response are valid | A tool is repeated or its result is interpreted incorrectly |
| Structured output | Schema validation and field semantics both pass | Valid JSON contains the wrong identifier or value |
| Interrupted or failed request | Retry and fallback logic leave a consistent state | A side effect is duplicated or partial state is abandoned |
| Routine successful task | Current quality and latency are preserved | An existing common success becomes a failure or times out |
Use only features documented for the candidate route. Mark an unsupported or unverified feature as a migration blocker for the dependent workflow instead of silently removing that case from the results.

A worked replay case for an existing Fable 5.1 agent
TEST-17, and the user’s instruction to update that ticket rather than create another. This is an illustrative application scenario, not a report of model behavior.Replay it in three forms: a fresh request carrying the required state, the original conversation history, and a resumed session after a simulated timeout. Keep the tool responses deterministic for the first pass. Then run a separate sandbox test with realistic tool errors.
| Acceptance item | Evidence to retain | Failure response |
|---|---|---|
| The correct ticket is updated | Tool name, arguments and resulting sandbox record | Reject a wrong ID or unintended create action |
| The user’s approved fields stay intact | Before/after record diff | Reject an unauthorized field change |
| A retry does not repeat a committed action | Application action ID and tool execution log | Stop the case and inspect retry/state handling |
| The final answer matches what happened | Compare answer with recorded tool result | Reject a claimed success after a failed update |
| Fallback can continue safely | Recovered state and a replay on the retained route | Do not expand traffic until recovery works |
Store a compact record for each attempt: case ID, model route, client and prompt versions, history transformation, tool log, pass/fail reasons, charges and review time. Keep an untested result blank rather than marking it passed. If fresh requests pass but historical sessions fail, investigate the history boundary before changing prompts everywhere. If both fail the same expected-state check, inspect the request and tool contract first.
The fixture makes an upgrade claim falsifiable: a fluent answer cannot hide a duplicate ticket or corrupted state. Apply the same pattern to your own application’s consequential actions.
{
"case_id": "ticket-update-17",
"before": {
"ticket_count": 1,
"ticket": {"id": "TEST-17", "owner": "Mina", "priority": "normal", "status": "open"}
},
"instruction": "Set TEST-17 priority to high. Keep its owner and status. Do not create another ticket.",
"expected_after": {
"ticket_count": 1,
"ticket": {"id": "TEST-17", "owner": "Mina", "priority": "high", "status": "open"}
},
"allowed_changed_fields": ["ticket.priority"],
"forbidden_operations": ["create_ticket", "close_ticket"]
}For the fresh-state case, supply the current ticket and instruction. For the history case, retain the earlier approval, creation response and update instruction. For recovery, let the sandbox apply the update but hide its response to simulate a timeout after a committed write. The app must reconcile the existing record before repeating an operation; an idempotency key helps only if the tool's documented contract actually honors it.
normal; a correct priority fails if the owner changes; a second ticket fails even if the first was updated correctly. Inspect both final state and the execution log: an extra write followed by a compensating edit could hide behind a correct final snapshot. When a tool response is lost, the assistant must not present an unverified update as confirmed.Save the starting state, submitted transcript, tool-call arguments, resulting state and final reply together. If both models fail the committed-write recovery case, fix the application's recovery mechanism before attributing the issue to model quality. This is a reusable acceptance fixture you can run on today's Fable 5.1 integration; the 5.5 column remains untested until a verified route exists.
Diagnose history-continuation failures
When a fresh request passes but its historical counterpart fails, compare the messages delivered to the model: tool-call/result pairs, retained user decisions and any previously generated content. Preserve the original history and version each transformation so you can identify which change affected the outcome.
Do not assume a cache entry, conversation reference or provider-specific field can be transferred to another model. Check the documented scope first. If a workflow requires a history conversion, test the converted history against the same expected state and record what information was removed or summarized.
Report passed, failed and untested counts separately for fresh and continued sessions, with reasons. A passing fresh-session group does not clear existing conversations for migration.
Move from replay to a reversible rollout
EvoLink can provide a common gateway surface for model selection, while model-specific behavior still needs validation. Keep the current model, candidate and fallback configuration explicit; do not populate the candidate with a guessed Fable 5.5 ID.
Turn replay results into a go/no-go decision
Write the decision before the rollout, including who can stop it and which state must be preserved. A higher average pass rate is insufficient when the improvement comes with a new critical failure.
| Hypothetical replay outcome | Decision | Next action |
|---|---|---|
| Candidate passes more tasks but duplicates one external action | Do not promote | Fix or isolate the action boundary, then replay both configurations |
| Fresh requests pass; historical sessions fail | Do not migrate existing conversations | Diagnose history conversion; evaluate new-session traffic separately |
| Quality passes but latency exceeds your predeclared limit | Do not make it the universal default | Test whether a clearly identified slow-task queue can tolerate it |
| Only one task class improves within its budget | Consider a bounded task-specific rollout | Validate classification errors and the fallback for that class |
| Results pass, but the retained route cannot resume state | Hold expansion | Restore a viable recovery path and test it |
These are example decisions, not observed Fable 5.5 results. Set numerical limits from your application’s requirements rather than copying an arbitrary success percentage. Review a sufficient range of normal and failure traffic for the intended rollout scope; a small sample remains uncertain.
For rollback, name the previous route and configuration version, pause new candidate assignments, identify in-flight tasks, reconcile completed side effects, and resume only from a known state. Record the trigger and recovery outcome. A fallback that exists only in configuration has not yet demonstrated recoverability.
What would make the upgrade worthwhile?
If the candidate brings no material benefit, retaining Fable 5.1 is a valid outcome while its route remains available and suitable. If only one workflow improves, migrating that workflow can be more defensible than changing every default. Neither choice requires treating an unverified release as a deadline.
FAQ
Can I replace the Fable 5.1 model ID with a guessed Fable 5.5 ID?
No. The exact candidate identifier and supported request contract need confirmation. A page slug is not a usable API model ID.
Is Fable 5.5 backward compatible with Fable 5.1?
The checked sources did not establish backward compatibility. Verify request fields, responses, tools and stateful behavior for the exact route you intend to use.
Can I reuse existing conversation histories?
Prepare them for testing, but do not assume compatibility. Test fresh sessions and history continuation separately, checking that important state and decisions survive.
Will prompt caches transfer to the candidate?
Do not assume cache portability. Verify the scope and billing rules for the candidate route and include cache-related charges in the evaluation.
Should I change prompts while switching models?
Start with a frozen baseline. If candidate-specific tuning is necessary, version it as a separate configuration and evaluate on tasks that were not used for tuning.
Is a shadow test safe for tool-using agents?
Only if it cannot unintentionally repeat external actions and is appropriate for the data being sent. Use recorded responses or a sandbox for early replay and account for duplicate request costs.
Is reverting the model setting enough for rollback?
Not always. Verify the fallback route and handle in-flight tasks, conversation state and external side effects. A configuration rollback does not undo an already executed action.
Should I migrate before an official announcement?
Sources and scope
- Anthropic model overview: model identity and documentation check.
- Anthropic news: release evidence check.
- EvoLink Fable 5.1: current-model baseline reference.


