Claude Opus 5 Is Here: Near-Fable Intelligence at Half the Price
Anthropic released Claude Opus 5 on July 24, 2026, and the headline is not raw power, it is economics. Opus 5 comes close to the frontier intelligence of Claude Fable 5 at half the price, and it costs exactly the same as its predecessor Opus 4.8: $5 per million input tokens and $25 per million output. It is available today on every platform as claude-opus-5, and it becomes the new default model on Claude Max and the strongest model available on Claude Pro.
The benchmark story is unusually strong. On Frontier-Bench v0.1, Opus 5 surpasses every other model and more than doubles Opus 4.8's performance at a lower cost per task. On CursorBench 3.2 at max effort, it lands within 0.5% of Fable 5's peak score at half the cost per task. On ARC-AGI 3, which tests solving genuinely novel problems, its score is three times the next-best model. On Zapier's AutomationBench its pass rate is around 1.5x the next-best model at the same cost, and even at its lowest effort setting it passes more tasks than any other model. On OSWorld 2.0 for computer use, it beats Fable 5's best result at just over a third of the cost.
What users report is a change in character more than capability. Opus 5 verifies its own work, catches its own logical faults during planning rather than after, and recovers from errors without you stepping in. On one benchmark task it was handed a drawing it had no way to actually view, so it wrote its own computer vision pipeline to extract the geometry from raw pixels and rebuilt the part, something no competing model managed in five attempts. Lovable reported it ahead of every model in its family on internal evals, up 22% over Opus 4.7 on their hardest agentic coding tasks, with far less run-to-run variance.
It is also Anthropic's most aligned model to date, scoring lowest of any recent model on misaligned behavior, with the lowest rates of deception. And unlike Fable 5, it carries no data retention requirement for general access.
Follow for more:
Course Registration: https://halaqa.app/enrollment?course=start-with-ai
The Release That Competes on Price, Not Just Power
Claude Opus 5 arrived on July 24, 2026, and Anthropic framed it around a claim about economics rather than raw capability: it comes close to the frontier intelligence of Fable 5 at half the price. Crucially, it costs exactly what Opus 4.8 cost - five dollars per million input tokens and twenty-five per million output - which means the upgrade is free in practical terms. It is available on every platform today as claude-opus-5, and it becomes the new default model on Claude Max as well as the strongest model accessible on Claude Pro. The timing explains the positioning. This is the fourth model Anthropic has shipped in under two months, following Mythos 5, Fable 5, and Sonnet 5, and Fable 5 drew real criticism from businesses over its token burn rate, with users blowing through budgets on tasks. Opus 5 is the direct answer to that: Anthropic describes it as designed to be used every day because it works more efficiently than other models. The competitive frame has shifted from who has the smartest model to who has the smartest model you can actually afford to run all day.
The Benchmarks: Doubling Its Predecessor, Tripling Its Rivals
The numbers are where the claim earns its weight. On Frontier-Bench v0.1, which evaluates valuable software engineering work, Opus 5 surpasses every other model and more than doubles Opus 4.8's performance while costing less per task. On CursorBench 3.2 at max effort, it comes within half a percentage point of Fable 5's peak score at half the cost per task, and it delivers more performance per dollar than any other model at high, xhigh, and max effort. On ARC-AGI 3, a test of solving genuinely novel problems the model has never seen, its score is three times the next-best model. On Zapier's AutomationBench, which measures completing real business tasks start to finish, its pass rate is about one and a half times the next-best model at the same cost per task - and even at its lowest effort setting it passes more tasks than any other model does at any setting. Zapier's CEO noted it took a raw account-health workbook and ran a full churn-prevention sequence end to end, hitting one hundred percent where previous models did not pass at all. On OSWorld 2.0 for computer use, it beats Fable 5's best result at just over a third of the cost. Anthropic states plainly that Opus 5 is the new state of the art on coding and knowledge work evaluations, while being honest that it remains behind Mythos 5 on cybersecurity tasks.
What Changed in Practice: Judgment, Self-Verification, and Fewer Turns
The most interesting reports about Opus 5 are not about scores; they describe a shift in behavior. Anthropic says it is much stronger at verifying its own work and iterating carefully until it succeeds, and three examples make that concrete. On one Frontier-Bench task it was given a drawing of a machine part and asked to rebuild it as a 3D model - but was deliberately given no way to view the drawing. It responded by writing its own computer vision pipeline to pull the geometry from raw pixels, then reconstructed the part, repeatedly, where no competing model succeeded in five attempts. Given a real bug in a popular open-source package manager, it found the root cause and fixed an edge case the community's own patch had missed, while a competing model fixed only the surface symptom and declared it resolved. An engineer at a trading firm had it build a market data feed for a new exchange in a single session; finding no live feed to validate against, it built its own test harness. Customers describe the same pattern from different angles: a frontend testing company reported it opened its pages in a browser at desktop and phone widths and caught a product hidden below the mobile fold before handing work back. A financial modeling team measured nine percentage points higher accuracy with a third fewer turns and tool calls and sixty percent less time. A legal team got comparable results while generating twenty-six percent fewer tokens than Opus 4.8 at max reasoning. And one engineer described it pushing back on a proposed design and not folding when pressed - instead narrowing its objection to a single question and proposing a compromise. Lovable, whose platform serves millions of builders, reported it ahead of every model in its family on internal evals, up twenty-two percent over Opus 4.7 on their hardest agentic coding tasks, with far less variance run to run.
Effort Levels, Science Gains, and the Safeguards You Will Notice
Three practical layers round out the release. First, effort control: Opus 5 exposes settings from low through medium, high, xhigh, and max, letting you optimize for intelligence or conserve tokens for faster and cheaper results. What matters is that it holds quality at lower effort better than its predecessor did - which is precisely why teams report comparable output at meaningfully lower token counts. The practical habit is to start at medium or high and raise the effort only when you actually hit a wall. Second, science: Opus 5 beats Opus 4.8 on every life sciences evaluation Anthropic ran, with the largest gains in organic chemistry, scoring ten point two percentage points higher on inferring molecular structures from spectroscopy data, and seven point seven points higher on predicting how protein sequence variations affect function. That makes it Anthropic's most capable generally available model for scientific research. Third, safeguards, which is where daily experience differs sharply from Fable 5. Opus 5's cyber classifiers are proportionally less restrictive and are expected to intervene around eighty-five percent less often than Fable 5's. It is allowed to find vulnerabilities in source code, while binary-based scanning, penetration testing, and exploit generation remain blocked; flagged requests fall back to Opus 4.8 by default across Claude.ai, Claude Code, and Cowork. Biology requests that get blocked on Fable 5 now route to Opus 5 rather than 4.8. On safety, Anthropic reports that Opus 5 does not advance the frontier of dual-use capability: it comes close to Mythos 5 at finding vulnerabilities but stays substantially behind at turning them into working exploits. And on alignment, its automated behavioral audit makes it Anthropic's most aligned model to date, scoring lowest on misaligned behavior with the lowest rates of deception and the least susceptibility to being tricked into misuse. Two beta updates shipped alongside it: mid-conversation tool changes that no longer invalidate the prompt cache, and automatic fallbacks on the API so flagged requests route to another model instead of failing outright. Unlike Fable 5, Opus 5 carries no data retention requirement for general access.
وصل Claude Opus 5: ذكاءٌ يقارب Fable بنصف الثمن
أطلقت Anthropic نموذج Claude Opus 5 في الرابع والعشرين من يوليو 2026، والعنوان هنا ليس القوّة الخام، بل الاقتصاد. فـ Opus 5 يقارب ذكاء Fable 5 المتقدّم بنصف ثمنه، ويكلّف تماماً ما كان يكلّفه سلفه Opus 4.8: خمسة دولارات لكل مليون رمز إدخال، وخمسة وعشرون لكل مليون إخراج. وهو متاح اليوم على جميع المنصّات باسم claude-opus-5، وصار النموذج الافتراضي الجديد في Claude Max، وأقوى نموذج متاح في Claude Pro.
وحكاية المعايير قويةٌ على غير المعتاد. فعلى Frontier-Bench v0.1، يتجاوز Opus 5 كل النماذج الأخرى، ويضاعف أداء Opus 4.8 بأكثر من مرّتين، وبكلفةٍ أقل للمهمة الواحدة. وعلى CursorBench 3.2 عند أقصى جهد، يحطّ على بُعد نصف نقطة مئوية من ذروة Fable 5، وبنصف الكلفة للمهمة. وعلى ARC-AGI 3، الذي يختبر حلّ مسائل جديدة كلياً، تبلغ نتيجته ثلاثة أضعاف أقرب منافسيه. وعلى AutomationBench من Zapier، يبلغ معدّل نجاحه نحو مرّة ونصف أقرب النماذج إليه بالكلفة ذاتها، بل إنه حتى عند أدنى إعداد جهدٍ لديه يجتاز مهامّ أكثر من أي نموذج آخر. وعلى OSWorld 2.0 لاستخدام الحاسوب، يتفوّق على أفضل نتيجة لـ Fable 5 بما يزيد قليلاً عن ثلث الكلفة.
أمّا ما يرويه المستخدمون فهو تبدّل في الطبع أكثر منه في القدرة. فـ Opus 5 يتحقّق من عمله بنفسه، ويلتقط أخطاءه المنطقية أثناء التخطيط لا بعد وقوعها، ويتعافى من العثرات دون أن تتدخّل. في إحدى مهامّ التقييم، سُلّم رسماً هندسياً لا يملك أي وسيلة لرؤيته فعلياً، فكتب بنفسه خطّ معالجةٍ للرؤية الحاسوبية يستخرج الأبعاد من البكسلات الخام، ثم أعاد بناء القطعة، وهو ما عجز عنه كل نموذج منافس عبر خمس محاولات. وأفادت Lovable بأنه تفوّق على كل نماذج عائلته في تقييماتها الداخلية، بزيادة 22% على Opus 4.7 في أصعب مهامّ البرمجة الوكيلة، وبتذبذبٍ أقلّ بكثير بين تشغيلةٍ وأخرى.
وهو كذلك أكثر نماذج Anthropic مواءمةً حتى اليوم، إذ سجّل أدنى معدّل سلوكٍ منحرف بين نماذجها الحديثة، وأقلّ نسب الخداع. وخلافاً لـ Fable 5، لا يفرض أي احتجازٍ إلزامي للبيانات في الوصول العام.
تابع حسابي على الانستغرام:
رابط الانضمام للدورة: https://halaqa.app/enrollment?course=start-with-ai
الخطوات
إصدارٌ يُنافس على السعر لا على القوّة وحدها
نزل Claude Opus 5 في الرابع والعشرين من يوليو 2026، وأطّرته Anthropic حول دعوى تخصّ الاقتصاد لا القدرة الخام: أنه يقارب ذكاء Fable 5 المتقدّم بنصف الثمن. والأهم أنه يكلّف تماماً ما كان يكلّفه Opus 4.8 - خمسة دولارات لكل مليون رمز إدخال، وخمسة وعشرون لكل مليون إخراج - ما يعني أنّ الترقية مجانيةٌ عملياً. وهو متاح على كل المنصّات اليوم باسم claude-opus-5، وصار النموذج الافتراضي الجديد في Claude Max، وأقوى نموذجٍ يمكن الوصول إليه في Claude Pro. والتوقيت يشرح الموقع. فهذا رابع نموذجٍ تطرحه Anthropic في أقلّ من شهرين، بعد Mythos 5 وFable 5 وSonnet 5، وقد واجه Fable 5 انتقاداً حقيقياً من الشركات بسبب معدّل استهلاكه للرموز، إذ كان المستخدمون يستنفدون ميزانياتهم على المهامّ. وOpus 5 هو الردّ المباشر على ذلك: تصفه Anthropic بأنه مُصمَّمٌ للاستخدام اليومي لأنه يعمل بكفاءةٍ أعلى من غيره. لقد انتقل إطار المنافسة من سؤال مَن يملك أذكى نموذج، إلى سؤال مَن يملك أذكى نموذجٍ تستطيع فعلاً أن تُشغّله طوال اليوم.
المعايير: يُضاعف سلفه ويُثلّث منافسيه
عند الأرقام تكتسب الدعوى ثقلها. فعلى Frontier-Bench v0.1، الذي يُقيّم أعمال هندسة البرمجيات ذات القيمة، يتجاوز Opus 5 كل النماذج ويُضاعف أداء Opus 4.8 بأكثر من مرّتين بكلفةٍ أقلّ للمهمة. وعلى CursorBench 3.2 عند أقصى جهد، يقترب إلى نصف نقطةٍ مئوية من ذروة Fable 5 بنصف الكلفة، ويُقدّم أداءً أعلى لكل دولار من أي نموذجٍ آخر عند مستويات الجهد العالية. وعلى ARC-AGI 3، وهو اختبارٌ لحلّ مسائل جديدة لم يرها النموذج قط، تبلغ نتيجته ثلاثة أضعاف أقرب منافسيه. وعلى AutomationBench من Zapier، الذي يقيس إنجاز مهامّ عملٍ حقيقية من أولها إلى آخرها، يبلغ معدّل نجاحه نحو مرّةٍ ونصف أقرب النماذج بالكلفة ذاتها - بل إنّ أدنى إعدادات جهده يجتاز مهامّ أكثر ممّا يجتازه أي نموذجٍ آخر عند أي إعداد. وذكر الرئيس التنفيذي لـ Zapier أنه أخذ جدول بياناتٍ خاماً لصحّة الحسابات ونفّذ تسلسل منع تسرّب العملاء كاملاً، فبلغ مئة بالمئة حيث لم تجتز النماذج السابقة أصلاً. وعلى OSWorld 2.0 لاستخدام الحاسوب، يتفوّق على أفضل نتائج Fable 5 بما يزيد قليلاً على ثلث الكلفة. وتقول Anthropic صراحةً إنّ Opus 5 هو طليعة المستوى الجديدة في تقييمات البرمجة وأعمال المعرفة، مع إقرارها بأنه يبقى خلف Mythos 5 في مهامّ الأمن السيبراني.
ما تغيّر عملياً: الحُكم، والتحقّق الذاتي، وعددُ جولاتٍ أقلّ
أطرف ما قيل عن Opus 5 ليس متعلّقاً بالدرجات، بل يصف تبدّلاً في السلوك. تقول Anthropic إنه أقوى بكثير في التحقّق من عمله والتكرار بعناية حتى ينجح، وثلاثة أمثلة تُجسّد ذلك. في إحدى مهامّ Frontier-Bench، سُلّم رسماً لقطعة آلة وطُلب منه إعادة بنائها نموذجاً ثلاثي الأبعاد - لكنّه حُرم عمداً من أي وسيلة لرؤية الرسم. فما كان منه إلا أن كتب خطّ معالجةٍ خاصاً به للرؤية الحاسوبية يستخرج الأبعاد من البكسلات الخام، ثم أعاد بناء القطعة، وكرّر ذلك بنجاح، حيث أخفق كل نموذجٍ منافس عبر خمس محاولات. وحين وُوجه بعلّةٍ حقيقية في مدير حزمٍ مفتوح المصدر واسع الانتشار، بلغ السبب الجذري وأصلح حالةً طرفية فاتت رقعة المجتمع نفسها، بينما اكتفى نموذجٌ منافس بمعالجة العَرَض الظاهر ثم أعلن حلّ المشكلة. وطلب منه مهندسٌ في شركة تداولٍ بناء تدفّق بيانات سوقٍ لبورصةٍ جديدة في جلسةٍ واحدة؛ وإذ لم يجد تدفّقاً حياً يتحقّق منه، بنى بنفسه منصّة اختبارٍ خاصة. ويصف العملاء النمط ذاته من زوايا شتّى: شركةُ اختبار واجهاتٍ أفادت أنه فتح صفحاته في متصفّح بعرض سطح المكتب والهاتف، فالتقط منتجاً مختفياً أسفل حدّ الشاشة قبل أن يُسلّم العمل. وفريقٌ للنمذجة المالية قاس دقّةً أعلى بتسع نقاط مئوية، بثلث أقلّ من الجولات واستدعاءات الأدوات، وبستّين بالمئة أقلّ من الوقت. وفريقٌ قانوني بلغ نتائج مقاربة بتوليد رموزٍ أقلّ بستٍّ وعشرين بالمئة من Opus 4.8 عند أقصى استدلال. ووصف مهندسٌ كيف اعترض النموذج على تصميمٍ اقترحه ولم يتراجع تحت الإلحاح - بل حصر اعتراضه في سؤالٍ واحد واقترح حلّاً وسطاً. أمّا Lovable، ومنصّتها تخدم ملايين البُناة، فأفادت بتفوّقه على كل نماذج عائلته في تقييماتها الداخلية، بزيادة اثنين وعشرين بالمئة على Opus 4.7 في أصعب مهامّ البرمجة الوكيلة، وبتذبذبٍ أقلّ بكثير بين تشغيلةٍ وأخرى.
مستويات الجهد، ومكاسب العلوم، والضمانات التي ستلاحظها
ثلاث طبقاتٍ عملية تُتمّ صورة الإصدار. أولاها التحكّم بالجهد: يُتيح Opus 5 إعدادات تمتدّ من low إلى medium وhigh وxhigh وmax، فتُحسّن للذكاء أو تدّخر الرموز لنتائج أسرع وأرخص. والمهمّ أنه يحافظ على الجودة عند الجهد المنخفض أفضل ممّا فعل سلفه - وهذا بالضبط ما يفسّر تقارير الفرق عن مخرجاتٍ مكافئة برموزٍ أقلّ بوضوح. والعادة العملية أن تبدأ عند medium أو high، ولا ترفع الجهد إلا حين تصطدم بجدارٍ فعلي. وثانيتها العلوم: يتفوّق Opus 5 على Opus 4.8 في كل تقييمات علوم الحياة التي أجرتها Anthropic، وأكبر مكاسبه في الكيمياء العضوية، إذ سجّل أعلى بعشر نقاطٍ ونقطتين من عشرة في استنتاج البِنى الجزيئية من بيانات التحليل الطيفي، وأعلى بسبع نقاطٍ وسبع من عشرة في التنبّؤ بأثر اختلافات تسلسل البروتين على وظيفته. وهذا يجعله أقدر نماذج Anthropic المتاحة للعموم في البحث العلمي. وثالثتها الضمانات، وهنا تختلف التجربة اليومية اختلافاً حاداً عن Fable 5. فمصنّفات Opus 5 السيبرانية أقلّ تقييداً نسبياً، ويُتوقّع أن تتدخّل بنحو خمسةٍ وثمانين بالمئة أقلّ من نظيرتها في Fable 5. ويُسمح له بإيجاد الثغرات في الشِّفرة المصدرية، بينما يبقى المسح القائم على الملفات الثنائية واختبار الاختراق وتوليد أدوات الاستغلال محجوباً؛ والطلبات المُعلّمة تتحوّل افتراضياً إلى Opus 4.8 عبر Claude.ai وClaude Code وCowork. أمّا طلبات البيولوجيا التي تُحجب على Fable 5 فصارت تُوجَّه إلى Opus 5 بدل 4.8. وعلى صعيد الأمان، تُفيد Anthropic بأنّ Opus 5 لا يدفع حدود القدرات مزدوجة الاستخدام: فهو يقارب Mythos 5 في إيجاد الثغرات، لكنه يبقى متأخّراً كثيراً في تحويلها إلى أدوات استغلالٍ عاملة. وعلى صعيد المواءمة، جعله التدقيق السلوكي الآلي أكثر نماذجها مواءمةً حتى اليوم، بأدنى درجةٍ في السلوك المنحرف، وأقلّ نسب الخداع، وأضعف قابليةٍ للانخداع نحو سوء الاستخدام. وصدر معه تحديثان تجريبيان: تغيير الأدوات في منتصف المحادثة دون إبطال ذاكرة البرومبت المؤقتة، والتحوّل التلقائي في الـ API بحيث تُوجَّه الطلبات المُعلّمة إلى نموذجٍ آخر بدل أن تفشل. وخلافاً لـ Fable 5، لا يفرض Opus 5 أي احتجازٍ إلزامي للبيانات في الوصول العام.
Prompt
# CLAUDE OPUS 5: LAUNCHED JULY 24, 2026 # ─── THE ESSENTIALS ─── # API model string: claude-opus-5 # Price: $5 / M input · $25 / M output (SAME as Opus 4.8) # Positioning: comes close to Fable 5 intelligence at HALF the price # New DEFAULT model on Claude Max · strongest model on Claude Pro # Available today on all platforms # Fast mode: ~2.5x default speed, at 2x base price (or usage credits in Claude Code) # No data retention requirement for general access (unlike Fable 5) # ─── HEADLINE BENCHMARKS ─── # Frontier-Bench v0.1 → beats ALL other models; MORE THAN DOUBLES Opus 4.8 # at a lower cost per task # CursorBench 3.2 → at max effort, within 0.5% of Fable 5 peak, HALF the cost # ARC-AGI 3 → 3x the score of the next-best model (novel problem solving) # Zapier AutomationBench → ~1.5x next-best pass rate at same cost; # even LOWEST effort beats every other model # OSWorld 2.0 (computer use) → beats Fable 5 best result at ~1/3 the cost # GDPval-AA v2, HLE, DeepSearchQA → best + most cost-efficient # Still BEHIND Mythos 5 on cybersecurity # ─── EFFORT LEVELS (the cost/intelligence dial) ─── # low · medium · high · xhigh · max # Lower effort = fewer tokens, faster, cheaper # Higher effort = deeper reasoning, better on hard tasks # Key finding: Opus 5 holds quality at LOWER effort better than 4.8 # (one legal team: similar performance with 26% fewer tokens vs 4.8 at max) # ─── SCIENCE GAINS OVER OPUS 4.8 ─── # Better on EVERY life-sciences eval Anthropic ran # Organic chemistry (molecular structure from spectroscopy): +10.2 points # Protein sequence-variation effects: +7.7 points # Now Anthropic's most capable GENERALLY AVAILABLE model for scientific research # ─── SAFEGUARDS & FALLBACKS (what you will actually notice) ─── # Cyber classifiers are LESS restrictive than Fable 5, expected to intervene # about 85% LESS often # Allowed: finding vulnerabilities in source code # Blocked: binary-based vuln scanning, penetration testing, exploit generation # Flagged requests fall back to Opus 4.8 by default in Claude.ai / Code / Cowork # Biology requests blocked on Fable 5 now route to Opus 5 (not 4.8) # ─── TWO BETA UPDATES SHIPPED ALONGSIDE ─── # Mid-conversation tool changes → swap Claude's tools mid-conversation WITHOUT # invalidating the prompt cache (big token saver for agents) # Automatic fallbacks on API → flagged requests auto-route to another model # instead of being blocked # ─── WHICH MODEL TO USE NOW ─── # Everyday work, coding, agents, analysis → OPUS 5 (best value at the frontier) # Multi-day autonomous projects / absolute hardest tasks → FABLE 5 # High-volume simple work → SONNET 5 / HAIKU # Rule: start at Opus 5 medium/high effort; raise effort only when you hit a wall.