Claude Fable 5.1: low effort now matches what full effort used to cost you

Anthropic released Claude Fable 5.1 on September 1, 2026, three months after Fable 5. The headline price did not move. It is still 10 dollars per million input tokens and 50 per million output. What moved is everything you get for that money.

Start with the claim that matters most. Anthropic says running Fable 5.1 at low or medium effort produces results comparable to or better than Fable 5 at full effort, at a meaningfully lower cost. Their accuracy-versus-cost curves on Terminal-Bench-Science show 5.1 beating 5 at nearly every effort tier. In plain terms, the setting you would reach for to save money now performs like the setting you used to pay a premium for.

The cost side changed through caching. Cache read pricing dropped 75 percent, down to 0.25 dollars per million tokens. Because cache reads make up a large share of real token usage in agentic work that keeps referencing the same big context, Anthropic puts the practical saving at roughly 25 percent for typical workloads and up to about 45 percent for context-heavy agentic tasks. Those figures come from four weeks of actual customer usage in August 2026 across Claude Enterprise, Claude Code, and the API, rather than from a projection.

Then the benchmark that stops you. On Terminal-Bench-Science, Fable 5.1 scored 52.6 percent against 24.7 percent for Fable 5. That is more than double, in three months. Anthropic says 5.1 outperforms both Fable 5 and Opus 5 across the board, and there is a detail underneath those numbers worth knowing: the model was evaluated with production safeguards switched on, so any case where its own safety system handed the task to an Opus model was scored as zero rather than excluded. The published scores understate what the model can actually do.

The science results are where it stops reading like a version bump. Mythos 5.1 designed protein binders whose binding affinities came in 10 times higher than the best entries submitted to Adaptyv Bio protein design competitions on three targets, with a hit rate near 50 percent across 12 targets against a field norm of 10 to 15 percent. Fable 5.1 trained a neural network that produced a new elevation map covering about a third of the surface of Venus from 30-year-old Magellan radar data, resolving detail down to 2 to 3 kilometers where the previous map managed 10 to 20. And at Millennium, it found the cause of a rare recurring crash that the firm's engineers, and every other model tested, had failed to explain for years.

Model ID is claude-fable-5-1, with a 1 million token context window and 128,000 token maximum output, available now on the Claude API, AWS, Google Cloud, and Azure.

Follow for more:

  • https://www.instagram.com/ai.with.mo/
  • Course Registration: https://halaqa.app/enrollment?course=start-with-ai

    The price stayed the same and the ceiling moved

    Anthropic shipped Fable 5.1 on September 1, 2026, three months after Fable 5, and left the headline price untouched at 10 dollars per million input tokens and 50 per million output. The saving arrives through caching instead. Cache read pricing fell 75 percent to 0.25 dollars per million tokens, and because cache reads account for a large share of real token consumption in agentic work that keeps referring back to the same large context, the effect on an actual bill is substantial. Anthropic puts it at roughly 25 percent lower cost for typical workloads and as much as 45 percent for context-heavy agentic tasks, and those numbers came out of four weeks of real customer usage during August 2026 across Claude Enterprise, Claude Code, and the API rather than a modeled estimate. Cognition made the point concrete on launch day by moving its Opus 5 traffic in Devin over to Fable 5.1, saying the cache pricing is what finally made a Fable-class model economical for work it had previously reserved for the cheaper tier. When the top model becomes affordable for jobs you were deliberately sending downmarket, the pricing change matters more than any single benchmark.

    Low effort now does what full effort used to do

    Here is the claim in Anthropic's own framing, because the precise wording is worth keeping. Running Fable 5.1 at low or medium effort produces results comparable to or better than Fable 5 at full effort, at meaningfully lower cost. Anthropic illustrated this with accuracy-versus-cost curves on Terminal-Bench-Science showing 5.1 above 5 at nearly every effort tier. So the honest version is comparable or better rather than a flat guarantee of beating it, and that is still a large statement: the cheap setting on the new model now sits where the expensive setting on the old one did. What it changes in practice is the habit. If you have been reaching for high or max because that was what Fable 5 needed to do good work, the default assumption is now wrong, and you are paying for it on every call. Drop one tier and measure the output before you decide you need the old level. It also helps to know that the default is not the same everywhere: Fable 5.1 runs at high effort by default in Claude Code, but medium in Claude Cowork and on Claude.ai. Two people comparing results on the same task may be comparing different effort settings without realizing it.

    More than double on science, and the score is understated

    On Terminal-Bench-Science, Fable 5.1 scored 52.6 percent where Fable 5 scored 24.7. More than double, in a three month gap between releases. Anthropic says 5.1 outperforms both Fable 5 and Opus 5 across its benchmark set, measured alongside OpenAI's GPT-5.6 Sol. The methodological detail underneath is the part most coverage skips, and it cuts in the model's favor. Anthropic evaluated Fable 5.1 with production safeguards switched on, which means every case where the model's own safety system intervened and handed the task to an Opus-class model was scored as a zero rather than dropped from the sample. On any benchmark that brushes against cybersecurity or biology, the published number therefore reflects the safeguards as much as the model, and the raw capability sits somewhere above what the chart shows. That is an unusually conservative way to report your own scores. It also explains why the customer reports read stronger than the benchmark table: Jane Street said 5.1 solved more of its internal coding benchmarks than either Fable 5 or Opus 5 while staying easier to follow across long multi-step work, and engineers at Datadog, Shopify, MongoDB, and the SpaceX AI team pointed to better root-cause analysis and longer unattended runs.

    Venus, protein binders, and a crash nobody could explain

    Three results carry more weight than any benchmark row. In molecular design, Mythos 5.1 was given open-source protein design and folding tools and asked to produce high-affinity protein binders, a starting point for many drug types. Anthropic sent the designs to two outside organizations for lab validation. On three targets the binding affinities came back 10 times higher than the best entries submitted to Adaptyv Bio's protein design competitions, and across 12 targets the hit rate was close to 50 percent where the field typically runs 10 to 15. In planetary science, Fable 5.1 trained a neural network that produced a new high-resolution elevation map of roughly a third of Venus, built from 30-year-old radar data from NASA's Magellan mission plus an existing partial map, resolving detail at 2 to 3 kilometers against 10 to 20 previously, with heights up to 25 percent more accurate. Anthropic is releasing it under a Creative Commons license ahead of the VERITAS and EnVision missions. And in computational biology, Mythos 5.1 wrote custom GPU kernels that sped up seven open-source genomics and protein models by as much as 2.5 times on an H100 with identical outputs, cutting estimated GPU costs for genome-wide analyses by 30 to 60 percent, in days rather than the weeks a team of performance engineers would need. Then there is the anecdote that lands hardest with anyone who has debugged production systems: at the investment firm Millennium, Fable 5.1 found the cause of a rare recurring crash that the firm's own engineers, and every other model they tried, had failed to explain for years.

    Claude Fable 5.1: الجهد المنخفض صار يوازي ما كان يكلفك الجهد الكامل

    أطلقت Anthropic نموذج Claude Fable 5.1 في الأول من سبتمبر 2026، بعد ثلاثة أشهر من Fable 5. والسعر المعلن لم يتحرك. لا يزال 10 دولارات لكل مليون رمز إدخال، و50 لكل مليون إخراج. الذي تحرك هو كل ما تحصل عليه مقابل هذا المبلغ.

    ابدأ من الادعاء الأهم. تقول Anthropic إن تشغيل Fable 5.1 على جهد منخفض أو متوسط يعطي نتائج مماثلة أو أفضل من Fable 5 على جهده الكامل، وبكلفة أدنى بوضوح. ومنحنيات الدقة مقابل الكلفة التي نشرتها على Terminal-Bench-Science تُظهر 5.1 متفوقاً على 5 عند كل درجة جهد تقريباً. وبلغة مباشرة: الإعداد الذي كنت تلجأ إليه لتوفير المال صار يؤدي مثل الإعداد الذي كنت تدفع علاوة مقابله.

    أما جانب الكلفة فتغير عبر التخزين المؤقت. انخفض سعر قراءة الكاش بنسبة 75 بالمئة، إلى 0.25 دولار لكل مليون رمز. ولأن قراءات الكاش تشكل حصة كبيرة من الاستهلاك الفعلي في الأعمال الوكيلة التي تعود إلى السياق الضخم نفسه مراراً، تقدر Anthropic التوفير العملي بنحو 25 بالمئة للأحمال المعتادة، وبما يصل إلى 45 بالمئة تقريباً للمهام الوكيلة كثيفة السياق. وهذه الأرقام مستخرجة من أربعة أسابيع من استخدام عملاء فعلي في أغسطس 2026 عبر Claude Enterprise وClaude Code والـ API، لا من تقدير نظري.

    ثم يأتي الرقم الذي يوقفك. على Terminal-Bench-Science، سجل Fable 5.1 نسبة 52.6 بالمئة مقابل 24.7 بالمئة لـ Fable 5. أي أكثر من الضعف، خلال ثلاثة أشهر. وتقول Anthropic إن 5.1 يتفوق على Fable 5 وعلى Opus 5 على امتداد الخط، وتحت هذه الأرقام تفصيلة تستحق المعرفة: قُيِّم النموذج والضمانات الإنتاجية مفعّلة، فكل حالة سلّم فيها نظام الأمان المهمة إلى نموذج من فئة Opus حُسبت صفراً بدل أن تُستبعد. أي أن الدرجات المنشورة أقل مما يستطيعه النموذج فعلاً.

    ونتائج العلوم هي حيث يتوقف الأمر عن كونه ترقية رقمية. صمم Mythos 5.1 روابط بروتينية جاءت شدة ارتباطها أعلى بعشر مرات من أفضل المشاركات في مسابقات تصميم البروتين لدى Adaptyv Bio على ثلاثة أهداف، بمعدل نجاح قارب 50 بالمئة عبر 12 هدفاً، مقابل معدل سائد في المجال يتراوح بين 10 و15 بالمئة. ودرّب Fable 5.1 شبكة عصبية أنتجت خريطة ارتفاعات جديدة تغطي نحو ثلث سطح كوكب الزهرة، اعتماداً على بيانات رادار عمرها ثلاثون عاماً من مسبار Magellan، بدقة تصل إلى 2 و3 كيلومترات حيث كانت الخريطة السابقة تقف عند 10 إلى 20. وفي شركة Millennium، وجد سبب انهيار متكرر نادر عجز مهندسو الشركة، وكل نموذج آخر جُرِّب، عن تفسيره لسنوات.

    معرّف النموذج هو claude-fable-5-1، بنافذة سياق مليون رمز وحد إخراج أقصى 128,000 رمز، متاح الآن على Claude API وAWS وGoogle Cloud وAzure.

    تابع حسابي على الانستغرام:

  • https://www.instagram.com/ai.with.mo/
  • رابط الانضمام للدورة: https://halaqa.app/enrollment?course=start-with-ai

    الخطوات

    السعر بقي كما هو والسقف ارتفع

    أطلقت Anthropic نموذج Fable 5.1 في الأول من سبتمبر 2026، بعد ثلاثة أشهر من Fable 5، وتركت السعر المعلن دون مساس عند 10 دولارات لكل مليون رمز إدخال و50 لكل مليون إخراج. أما التوفير فيأتي من التخزين المؤقت. انخفض سعر قراءة الكاش 75 بالمئة إلى 0.25 دولار لكل مليون رمز، ولأن قراءات الكاش تستحوذ على حصة كبيرة من الاستهلاك الفعلي في الأعمال الوكيلة التي تعود مراراً إلى السياق الضخم نفسه، فإن الأثر على الفاتورة الحقيقية كبير. تقدر Anthropic ذلك بنحو 25 بالمئة أقل للأحمال المعتادة، وبما يبلغ 45 بالمئة للمهام الوكيلة كثيفة السياق، وهذه الأرقام خرجت من أربعة أسابيع من استخدام عملاء فعلي خلال أغسطس 2026 عبر Claude Enterprise وClaude Code والـ API، لا من تقدير نظري. وقد جسّدت Cognition الأمر يوم الإطلاق حين نقلت حركة Opus 5 في Devin إلى Fable 5.1، قائلة إن تسعير الكاش هو ما جعل أخيراً نموذجاً من فئة Fable مجدياً اقتصادياً لأعمال كانت تحيلها سابقاً إلى الفئة الأرخص. وحين يصبح النموذج الأعلى في متناول أعمال كنت تتعمد إرسالها إلى الأسفل، فإن تغيير التسعير يزن أكثر من أي معيار منفرد.

    الجهد المنخفض صار يفعل ما كان يفعله الجهد الكامل

    هذا هو الادعاء بصياغة Anthropic نفسها، لأن الدقة في اللفظ تستحق الحفاظ عليها. تشغيل Fable 5.1 على جهد منخفض أو متوسط يعطي نتائج مماثلة أو أفضل من Fable 5 على جهده الكامل، وبكلفة أدنى بوضوح. وقد وضّحت Anthropic ذلك بمنحنيات الدقة مقابل الكلفة على Terminal-Bench-Science، تُظهر 5.1 فوق 5 عند كل درجة جهد تقريباً. فالصيغة الأمينة هي مماثل أو أفضل، لا ضمان قاطع بالتفوق، وهذا وحده قول كبير: الإعداد الرخيص في النموذج الجديد صار حيث كان الإعداد الغالي في القديم. والذي يتغير عملياً هو العادة. فإن كنت تلجأ إلى الجهد العالي أو الأقصى لأن ذلك ما كان يحتاجه Fable 5 ليُخرج عملاً جيداً، فافتراضك الافتراضي صار خاطئاً، وأنت تدفع ثمنه في كل نداء. انزل درجة واحدة وقِس المخرجات قبل أن تقرر أنك تحتاج المستوى القديم. ومن المفيد أن تعرف أيضاً أن الإعداد الافتراضي ليس واحداً في كل مكان: يعمل Fable 5.1 على جهد عالٍ افتراضياً في Claude Code، لكن على متوسط في Claude Cowork وعلى Claude.ai. وقد يقارن شخصان نتائج المهمة نفسها وهما في الحقيقة يقارنان درجتي جهد مختلفتين دون أن ينتبها.

    أكثر من الضعف في العلوم، والدرجة أقل من الواقع

    على Terminal-Bench-Science، سجل Fable 5.1 نسبة 52.6 بالمئة حيث سجل Fable 5 نسبة 24.7. أكثر من الضعف، في فجوة ثلاثة أشهر بين إصدارين. وتقول Anthropic إن 5.1 يتفوق على Fable 5 وعلى Opus 5 عبر مجموعة معاييرها، مقيسة إلى جانب GPT-5.6 Sol من OpenAI. أما التفصيلة المنهجية تحت ذلك فهي ما تتجاوزه معظم التغطيات، وهي تصب في مصلحة النموذج. قيّمت Anthropic النموذج والضمانات الإنتاجية مفعّلة، ما يعني أن كل حالة تدخّل فيها نظام الأمان الخاص به وسلّم المهمة إلى نموذج من فئة Opus حُسبت صفراً بدل أن تُسقط من العينة. وعلى أي معيار يلامس الأمن السيبراني أو البيولوجيا، فإن الرقم المنشور يعكس الضمانات بقدر ما يعكس النموذج، والقدرة الخام تقع في مكان أعلى مما يُظهره الرسم. وهذه طريقة متحفظة على غير المعتاد في عرض درجاتك الخاصة. وهي تفسر أيضاً لماذا تبدو تقارير العملاء أقوى من جدول المعايير: قالت Jane Street إن 5.1 حل من معاييرها البرمجية الداخلية أكثر مما حله Fable 5 أو Opus 5، مع بقائه أسهل في المتابعة عبر الأعمال الطويلة متعددة الخطوات، وأشار مهندسون في Datadog وShopify وMongoDB وفريق الذكاء الاصطناعي في SpaceX إلى تحليل أفضل للأسباب الجذرية وتشغيلات أطول دون إشراف.

    الزهرة، والروابط البروتينية، وانهيار عجز الجميع عن تفسيره

    ثلاث نتائج تزن أكثر من أي سطر في جدول معايير. في التصميم الجزيئي، أُعطي Mythos 5.1 أدوات مفتوحة المصدر لتصميم البروتينات وطيّها، وطُلب منه إنتاج روابط بروتينية عالية الألفة، وهي نقطة بداية لأنواع دوائية كثيرة. وأرسلت Anthropic التصاميم إلى منظمتين خارجيتين للتحقق المخبري. وعلى ثلاثة أهداف عادت شدة الارتباط أعلى بعشر مرات من أفضل المشاركات في مسابقات تصميم البروتين لدى Adaptyv Bio، وعبر 12 هدفاً بلغ معدل النجاح ما يقارب 50 بالمئة حيث يتراوح المعتاد في المجال بين 10 و15. وفي علوم الكواكب، درّب Fable 5.1 شبكة عصبية أنتجت خريطة ارتفاعات جديدة عالية الدقة لنحو ثلث كوكب الزهرة، مبنية على بيانات رادار عمرها ثلاثون عاماً من مهمة Magellan التابعة لناسا إضافة إلى خريطة جزئية قائمة، تحل التفاصيل عند 2 إلى 3 كيلومترات مقابل 10 إلى 20 سابقاً، بارتفاعات أدق بما يصل إلى 25 بالمئة. وتنشرها Anthropic برخصة المشاع الإبداعي قبيل مهمتي VERITAS وEnVision. وفي البيولوجيا الحاسوبية، كتب Mythos 5.1 نوى معالجة رسومية مخصصة سرّعت سبعة نماذج مفتوحة المصدر في الجينوم والبروتين بما يبلغ 2.5 ضعف على معالج H100 بمخرجات مطابقة، فخفضت الكلفة المقدرة للتحليلات الجينومية الواسعة بين 30 و60 بالمئة، في أيام بدل الأسابيع التي يحتاجها فريق من مهندسي الأداء. ثم تأتي الحكاية التي تصيب أعمق أثر لدى كل من صحّح أنظمة إنتاجية: في شركة الاستثمار Millennium، وجد Fable 5.1 سبب انهيار متكرر نادر عجز مهندسو الشركة أنفسهم، وكل نموذج آخر جرّبوه، عن تفسيره لسنوات.

    Prompt

    # CLAUDE FABLE 5.1, RELEASED SEPTEMBER 1 2026
    
    # THE ESSENTIALS
    # Model ID: claude-fable-5-1
    # Context: 1,000,000 tokens. Max output: 128,000 tokens.
    # Available: Claude API, AWS, Google Cloud, Microsoft Azure
    # Headline price UNCHANGED: 10 USD per M input, 50 USD per M output
    # Fable 5.1 and Mythos 5.1 are the SAME model with different safeguards.
    #   Mythos 5.1 is trusted-access only (cyber and life sciences programs).
    
    # THE COST CHANGE
    # Cache reads cut 75 percent, down to 0.25 USD per M tokens
    # Roughly 25 percent lower total cost on typical workloads
    # Up to roughly 45 percent lower on context-heavy agentic tasks
    # Based on 4 weeks of real customer usage in August 2026 across
    #   Claude Enterprise, Claude Code, and the API
    
    # THE EFFORT CLAIM (the one that matters for your bill)
    # Anthropic: Fable 5.1 at LOW or MEDIUM effort gives results comparable
    #   to or better than Fable 5 at FULL effort, at meaningfully lower cost.
    # Their accuracy-vs-cost curves show 5.1 above 5 at nearly every tier.
    #
    # Effort defaults differ by surface:
    #   Claude Code: High
    #   Claude Cowork: Medium
    #   Claude.ai: Medium
    # Practical habit: start one tier lower than you used to. Raise it only
    #   when you hit a wall on a specific task.
    
    # BENCHMARK HEADLINE
    # Terminal-Bench-Science: 52.6 percent, against 24.7 percent for Fable 5
    # Anthropic says 5.1 beats both Fable 5 and Opus 5 across the board
    # IMPORTANT: scored WITH production safeguards active. Safety handoffs to
    #   Opus models counted as ZERO rather than being excluded, so the
    #   published numbers understate raw capability.
    
    # WHAT CHANGED IN THE SAFEGUARDS
    # Cybersecurity: about 60 percent fewer benign requests flagged.
    #   Vulnerability DISCOVERY is now allowed as a defensive measure.
    #   Exploit development, penetration testing, and binary-based scanning
    #   are still redirected to Opus-class models.
    # Biology and medical: fallback rates cut about 85 percent for everyday
    #   questions such as interpreting lab results or general health queries.
    
    # SCIENCE RESULTS WORTH KNOWING
    # Protein design (Mythos 5.1): binding affinities 10x higher than the best
    #   entries in Adaptyv Bio competitions on 3 targets. Hit rate near 50
    #   percent across 12 targets, against a field norm of 10 to 15 percent.
    # Venus mapping (Fable 5.1): new elevation map of about a third of the
    #   planet from 30-year-old Magellan radar data. Detail down to 2 to 3 km
    #   versus 10 to 20 km before. Heights up to 25 percent more accurate.
    #   Released under a Creative Commons license.
    # GPU kernels (Mythos 5.1): sped up 7 open-source genomics and protein
    #   models by as much as 2.5x on an H100 with identical outputs. Cut
    #   estimated GPU cost for genome-wide analyses by 30 to 60 percent.
    #   Work that would take a team weeks, done in days.
    
    # WHAT EARLY CUSTOMERS REPORTED
    # Jane Street: solved more internal coding benchmarks than Fable 5 or Opus 5
    # Cognition: shifting Devin Opus 5 traffic to Fable 5.1 from day one,
    #   citing cache pricing as what finally made a Fable-class model economical
    # Every (Dan Shipper): roughly twice as fast as Opus 5, about half the tokens
    # Rogo: matched Fable 5 accuracy using 20 percent fewer tokens
    # Millennium: found a rare crash cause its engineers had chased for years
    
    # MIGRATION NOTES BEFORE YOU SWITCH
    # 1. Forced tool calls have a documented compatibility change. Test yours.
    # 2. New API accounts created from launch day can no longer manually edit
    #    Claude prior context while preserving the thinking transcript. This is
    #    an anti-distillation measure. Existing accounts are not affected yet.
    # 3. Enterprise Frontier Safeguards (customer-controlled data handling)
    #    rolls out in phases this fall. Zero data retention is available in the
    #    interim for eligible customers.
    # 4. Outputs from models released after August 2 2026 carry an invisible
    #    watermark under the EU AI Act Code of Practice.
    
    # HOW TO ACTUALLY SAVE MONEY WITH IT
    # Drop your effort setting one tier and measure the output before assuming
    #   you need the old level.
    # Structure long agentic work so it reuses the same cached context. That is
    #   where the 75 percent cache cut turns into a real bill difference.
    # Reserve high and max effort for the specific tasks that failed at medium.