I left Claude for ChatGPT. Opus 5.5 wins on paper, GPT-6 Sol wins on price
I canceled my Claude subscription because ChatGPT's models were finishing more of my work. That was a decision about results, and results cut both ways. On 2026-09-22 Anthropic released Opus 5.5 and OpenAI released GPT-6 Sol on the same day. Anthropic's own numbers now give me a reason to test Claude again.
Neither launch post tests the two new models against each other, but they share reference points. On AutomationBench both list Opus 5 at 26.9% and Fable 5.1 at 31.4%. Against those shared marks, Anthropic reports Opus 5.5 at 40.0% and OpenAI reports GPT-6 Sol at 33.2%. On paper, Opus 5.5 is ahead.
On price, Sol is ahead. It costs 2 dollars per million input tokens and 10 per million output tokens, half of Opus 5.5's 4 and 20, and OpenAI pairs it with higher usage limits in ChatGPT Work and Codex. Anthropic answers with raised five-hour limits and a usage reset subscribers can save for later. Which of those is cheaper for you depends on how many steps each model needs on your work, and no launch post can tell you that.
Going back costs me little, because my project history lives in plain files on my own machine. That is the job of the Memory Engine I built: it files sessions from ChatGPT Codex, Claude Code and three other assistants into one folder that any of them can read.
The copy block is the test I am running: one unfinished task and the same files for both tools, with four things recorded for each: whether the result passes its check, how many corrections it needs, how long it takes and how far the usage meter moves.
Follow for more:
Course Registration: https://halaqa.app/enrollment?course=start-with-ai
I left Claude because ChatGPT was finishing more of my work
I canceled Claude for a plain reason: the ChatGPT models were getting better results on the work I actually do. By mid-September, about 1.5 billion of the roughly 2.45 billion tokens I have used had gone through ChatGPT, against almost 200 million in Claude Code. The run that settled it was a 13-hour job in which 45 agents finished what I had budgeted three weeks for, planned by GPT-6 Astra and carried out by GPT-5.6 Sol. None of that was loyalty, and the same rule applies in the other direction. A subscription earns its place by the work it finishes, not by the company I happened to start with. Opus 5.5 gives me a concrete reason to apply that rule again, and the reason is in Anthropic's own numbers.
Both launch posts use the same yardsticks, and Opus 5.5 leads on them
On AutomationBench, a test of business workflows across apps, Anthropic's post and OpenAI's post print the same scores for Opus 5 (26.9%) and Fable 5.1 (31.4%), even though neither post tests Opus 5.5 against GPT-6 Sol directly. Against those shared marks, Anthropic reports Opus 5.5 at 40.0% and OpenAI reports GPT-6 Sol at 33.2% at xhigh effort. FrontierCode, which grades whether a coding agent's changes are ready to merge into a real codebase, points the same way less directly. Anthropic puts Opus 5.5 at 54.4% and Fable 5.1 at 50.3%. OpenAI's text says Sol matches Fable 5.1 at xhigh effort. Read this with its limits. Each company ran its own model in its own setup, and the effort settings are not the same. Anthropic's own table also puts GPT-6 Astra at 41.4% on AutomationBench, ahead of Opus 5.5.
GPT-6 Sol costs half as much, and that is its whole argument
At API rates, Opus 5.5 costs 4 dollars per million input tokens and 20 per million output tokens. GPT-6 Sol costs 2 and 10, after OpenAI cut its price by half from GPT-5.6 Sol. OpenAI also publishes a cost per task on AutomationBench: 27 cents for Sol at xhigh, with Opus 5 at 11.1 times that. Anthropic publishes no cost per task for Opus 5.5. The bill also depends on steps, because a model that needs fewer of them spends fewer tokens. That is exactly what Anthropic's early testers report: GitHub says Opus 5.5 solved more terminal tasks than Opus 5 in less than half the steps, and Lovable reports a third to half fewer steps. How much of the price gap that closes on your work is something only a run on your work can show. Subscribers pay a third bill, the usage meter. Anthropic raised five-hour limits on Pro, Max, Team and seat-based Enterprise, and gave subscription users a reset they can save and spend when they choose. OpenAI promises higher usage limits with Sol, which is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu, and not yet in the regular chat.
Going back is cheap for me because my context never lived in either app
The expensive part of switching is the first week, when the new model knows nothing about the project: the approach you already rejected and why, the client note that changed the brief, the rule you have given twice, the file it must never touch. I avoid that week because my project history lives in plain files on my own machine. The Memory Engine I built reads sessions from ChatGPT Codex, Claude Code, Claude Cowork, Gemini and Cursor every 15 minutes and files them into one folder, so Opus 5.5 can start from the decisions already in those files, including the ones I made in Codex. That is what makes a fair test possible. I am giving Opus 5.5 in Claude Code and GPT-6 Sol in Codex the same unfinished task from my own work, in the same folder, with the same files. For each one I will record whether the result passes its check without my help, how many corrections it needs before I would ship it, how long it takes and how far the usage meter moves. The model that finishes more accepted work per allowance gets the subscription. The copy block is that test, written so you can run it on your own task this week.
ألغيت اشتراكي في Claude، وOpus 5.5 يعيدني لأجرّبه: الأرقام في صفّه، والسعر في صفّ GPT-6 Sol
ألغيت اشتراكي في Claude ونقلت عملي إلى ChatGPT لسبب واحد: نماذج OpenAI كانت تعطيني نتائج أفضل في عملي الحقيقي. لا ولاء عندي لأي شركة، أدفع لمن يُنجز. ومن قرابة 2.45 مليار token استخدمتها حتى منتصف سبتمبر 2026، مرّ نحو 1.5 مليار عبر ChatGPT.
وفي يوم واحد أطلقت Anthropic نموذج Opus 5.5 وأطلقت OpenAI نموذج GPT-6 Sol. لا تقارن أيّ من الشركتين نموذجها بنموذج الأخرى، لكن المنشورين يذكران النتيجة نفسها لنموذجين أقدم من Claude على اختبار AutomationBench، الذي يقيس مهام عمل تتنقّل بين تطبيقات كثيرة، وهذا يكفي لوضع الرقمين جنبًا إلى جنب: Opus 5.5 يسجّل 40.0% في منشور Anthropic، وGPT-6 Sol يسجّل 33.2% في منشور OpenAI.
الأرقام إذن تميل إلى Opus 5.5. أما السعر فيميل بوضوح إلى Sol: في API يدفع المطوّر دولارين لكل مليون token من المدخلات بدل 4 دولارات مع Opus 5.5، و10 دولارات للمخرجات بدل 20. أي نصف السعر بالضبط.
ولا يحسم هذا قراري. كل شركة اختبرت نموذجها في ظروفها هي، ومستويات الجهد مختلفة، وجدول Anthropic نفسه يضع GPT-6 Astra، النموذج الأعلى عند OpenAI، فوق Opus 5.5. وما يهمني في النهاية سؤال آخر: كم مهمة من مهامي تنتهي مقبولة من دون أن أتدخل؟
لذلك سأجرّب Opus 5.5 قبل أن أقرر أي اشتراك أحتفظ به: أعطيه وأعطي GPT-6 Sol المهمة نفسها، وأحسب ما يُنجزه كل منهما مقابل ما يستهلكه من حصة الاشتراك. وسأحكم بالعمل الذي يُسلَّم فعلًا. والانتقال بين الأداتين لا يكلّفني أسبوعًا ضائعًا، لأن تاريخ مشاريعي لا يسكن داخل أي منهما.
تابعني على Instagram:
التسجيل في الدورة: https://halaqa.app/enrollment?course=start-with-ai
الخطوات
خرجتُ من Claude لأن نتائج ChatGPT في عملي كانت أفضل
حتى منتصف سبتمبر 2026 استخدمت قرابة 2.45 مليار token عبر أنظمة الذكاء الاصطناعي كلها. نحو 1.5 مليار منها مرّ عبر ChatGPT، وقرابة 200 مليون فقط في Claude Code. هذه أرقام من لوحات الاستهلاك عندي، وتفصيلها في أين ذهبت 2.45 مليار token. وأكبر تجربة خضتها هناك كانت يوم عمل امتد 13 ساعة، أنهى فيه 45 agent عملًا كنت قدّرته بثلاثة أسابيع. خطّطت له مع GPT-6 Astra ونفّذته مع GPT-5.6 Sol، وتجد تفاصيل تلك التجربة هنا. لم تكن في القرار عاطفة. الاشتراك يذهب إلى الأداة التي تُنهي عملي، وبالمعيار نفسه أعود الآن لأختبر Claude بعد صدور Opus 5.5.
منشورا الإطلاق يسمحان بمقارنة لم تُجرِها أيّ من الشركتين
لا تضع Anthropic نموذجها أمام GPT-6 Sol، ولا تضع OpenAI نموذجها أمام Opus 5.5. لكن الجدولين يلتقيان عند AutomationBench، وهو اختبار لمهام عمل تتنقّل بين تطبيقات كثيرة: كلاهما يعطي Opus 5 نسبة 26.9%، وكلاهما يعطي Fable 5.1 نسبة 31.4%. وما دام الأساس واحدًا، يمكن أن تقرأ الجدولين معًا. على هذا الأساس تنشر Anthropic لنموذج Opus 5.5 نسبة 40.0%، وتنشر OpenAI لنموذج GPT-6 Sol نسبة 33.2% عند مستوى الجهد xhigh. وفي FrontierCode، وهو اختبار يسأل هل تصلح تعديلات الكود للدمج في مشروع حقيقي، يسجّل Opus 5.5 نسبة 54.4% مقابل 50.3% لنموذج Fable 5.1، بينما تكتفي OpenAI بالقول إن Sol يعادل Fable 5.1 عند xhigh، من غير أن تنشر رقمه. ولهذه القراءة حدود لا أتجاهلها. كل شركة شغّلت نموذجها في بيئتها الخاصة، ومستوى الجهد ليس واحدًا في الحالتين. والأهم أن جدول Anthropic نفسه يضع GPT-6 Astra، أقوى نماذج OpenAI، عند 41.4%، أي فوق Opus 5.5.
GPT-6 Sol بنصف السعر، لكن عدد الخطوات يدخل في الفاتورة أيضًا
في API يكلّف كل مليون token من المدخلات مع Opus 5.5 أربعة دولارات، وكل مليون من المخرجات 20 دولارًا. أما GPT-6 Sol فبدولارين و10 دولارات، بعد أن خفّضت OpenAI سعره 50% عن GPT-5.6 Sol. وتنشر OpenAI كلفة المهمة الواحدة على AutomationBench: 0.27 دولار مع Sol عند xhigh، و11.1 ضعف ذلك مع Opus 5. ولا تنشر Anthropic رقمًا مقابلًا لنموذج Opus 5.5. لكن سعر الـ token لا يصنع الفاتورة وحده، فالنموذج الذي يصل إلى الحل بخطوات أقل يقرأ ويكتب أقل. تنقل Anthropic عن مختبريها الأوائل أن GitHub رأت Opus 5.5 يحل مهام terminal أكثر مما يحله Opus 5، في أقل من نصف الخطوات، وأن Lovable رأت الخطوات تقل بما بين الثلث والنصف. هذه المقارنة مع Opus 5 وليست مع Sol، فلا أعرف بعد إن كانت تكفي لسد فرق السعر. وأغلب من يقرأ هذا لا يدفع عبر API بل باشتراك، والعدّاد هناك مختلف. رفعت Anthropic حدود الاستخدام المحسوبة على كل خمس ساعات لباقات Pro، Max، Team وEnterprise بنظام المقاعد، ومنحت المشتركين reset لحدود الاستخدام يحتفظون به ويستعملونه متى شاؤوا. وتعد OpenAI بحدود أعلى مع Sol، وهو متاح في ChatGPT Work وCodex لمشتركي Plus، Pro، Business، Enterprise وEdu، ولم يصل بعد إلى المحادثة العادية.
يبقى الاشتراك للنموذج الذي يُنهي عملًا مقبولًا أكثر بالحصة نفسها
الاختبار الذي سأجريه بسيط. آخذ مهمة حقيقية لم تكتمل بعد، وأسلّمها إلى Opus 5.5 داخل Claude Code وإلى GPT-6 Sol داخل Codex، في المجلد نفسه ومع الملفات نفسها. ثم أسجّل لكل منهما ما اجتاز الفحص من عمله دون مساعدة، وعدد التصحيحات التي احتاجتها النتيجة، والوقت الذي استغرقه، والقدر الذي استهلكه من عدّاد الاشتراك. النموذج الذي يُنهي عملًا مقبولًا أكثر بالحصة نفسها يحتفظ بالاشتراك. التعليمات جاهزة في مربع النسخ أسفل الصفحة، ويمكنك تشغيلها على مهمة من مهامك. أغلى ما في تغيير الأداة هو الأسبوع الأول، حين لا يعرف النموذج الجديد شيئًا عن مشروعك: الحلول التي جرّبتها ورفضتها ولماذا، وملاحظة العميل التي غيّرت المطلوب، والقواعد التي كرّرتها على النموذج السابق مرة بعد مرة، والملفات التي طلبت منه ألا يلمسها. أنا لا أدفع هذه الكلفة، لأن تاريخ مشاريعي محفوظ في ملفات نصية عادية على جهازي. بنيت محرّك الذاكرة لهذا الغرض: كل 15 دقيقة يقرأ الجلسات من ChatGPT Codex، Claude Code، Claude Cowork، Gemini وCursor، ويحفظها نصًا عاديًا في مجلد واحد على الجهاز، فيقرؤه أيّ منها بعد ذلك. فحين أفتح Claude Code للاختبار، يجد Opus 5.5 تاريخ المهمة أمامه من اليوم الأول.
Prompt
# OPUS 5.5 VS GPT-6 SOL: THE ONE-TASK TEST # Run this before you renew, upgrade or cancel either subscription. # --- PICK THE TASK --- # One real, unfinished task from your own work, with a check you trust: # a test suite, or a brief a client would sign off. # Pick something you would otherwise do yourself this week. # --- GIVE BOTH THE SAME START --- # Put the context in files both tools read: # AGENTS.md for your standing rules and how finished work is checked. # STATE.md for where the work is, what was decided and why, what was # rejected, and what comes next. # Codex reads AGENTS.md by default. Claude Code reads it when the folder has # no CLAUDE.md; otherwise put "@AGENTS.md" on the first line of CLAUDE.md. # # Open Claude Code on Opus 5.5 and Codex on GPT-6 Sol, each on its own branch # or copy of the folder. Send both the same first message: # # Read AGENTS.md, then STATE.md. Finish the task under "Next" in STATE.md. # Run the checks named there before you tell me you are done. When you finish, # rewrite STATE.md so another assistant could continue without asking me. # --- RECORD FOR EACH --- # 1. Did it pass the check without my help? # 2. How many corrections before I would ship it? # 3. Wall time from the first message to an accepted result. # 4. The usage meter before and after the run. # --- DECIDE --- # Keep the tool that finishes more accepted work per allowance on the task # you do most often. A close result is a tie: run a second task of a # different kind. Run the test again when either company changes a model # or a limit.