I left Claude for ChatGPT. Opus 5.5 wins on paper, GPT-6 Sol wins on price

I canceled my Claude subscription because ChatGPT's models were finishing more of my work. That was a decision about results, and results cut both ways. On 2026-09-22 Anthropic released Opus 5.5 and OpenAI released GPT-6 Sol on the same day. Anthropic's own numbers now give me a reason to test Claude again.

Neither launch post tests the two new models against each other, but they share reference points. On AutomationBench both list Opus 5 at 26.9% and Fable 5.1 at 31.4%. Against those shared marks, Anthropic reports Opus 5.5 at 40.0% and OpenAI reports GPT-6 Sol at 33.2%. On paper, Opus 5.5 is ahead.

On price, Sol is ahead. It costs 2 dollars per million input tokens and 10 per million output tokens, half of Opus 5.5's 4 and 20, and OpenAI pairs it with higher usage limits in ChatGPT Work and Codex. Anthropic answers with raised five-hour limits and a usage reset subscribers can save for later. Which of those is cheaper for you depends on how many steps each model needs on your work, and no launch post can tell you that.

Going back costs me little, because my project history lives in plain files on my own machine. That is the job of the Memory Engine I built: it files sessions from ChatGPT Codex, Claude Code and three other assistants into one folder that any of them can read.

The copy block is the test I am running: one unfinished task and the same files for both tools, with four things recorded for each: whether the result passes its check, how many corrections it needs, how long it takes and how far the usage meter moves.

Follow for more:

  • https://www.instagram.com/ai.with.mo/
  • Course Registration: https://halaqa.app/enrollment?course=start-with-ai

    I left Claude because ChatGPT was finishing more of my work

    I canceled Claude for a plain reason: the ChatGPT models were getting better results on the work I actually do. By mid-September, about 1.5 billion of the roughly 2.45 billion tokens I have used had gone through ChatGPT, against almost 200 million in Claude Code. The run that settled it was a 13-hour job in which 45 agents finished what I had budgeted three weeks for, planned by GPT-6 Astra and carried out by GPT-5.6 Sol. None of that was loyalty, and the same rule applies in the other direction. A subscription earns its place by the work it finishes, not by the company I happened to start with. Opus 5.5 gives me a concrete reason to apply that rule again, and the reason is in Anthropic's own numbers.

    Both launch posts use the same yardsticks, and Opus 5.5 leads on them

    On AutomationBench, a test of business workflows across apps, Anthropic's post and OpenAI's post print the same scores for Opus 5 (26.9%) and Fable 5.1 (31.4%), even though neither post tests Opus 5.5 against GPT-6 Sol directly. Against those shared marks, Anthropic reports Opus 5.5 at 40.0% and OpenAI reports GPT-6 Sol at 33.2% at xhigh effort. FrontierCode, which grades whether a coding agent's changes are ready to merge into a real codebase, points the same way less directly. Anthropic puts Opus 5.5 at 54.4% and Fable 5.1 at 50.3%. OpenAI's text says Sol matches Fable 5.1 at xhigh effort. Read this with its limits. Each company ran its own model in its own setup, and the effort settings are not the same. Anthropic's own table also puts GPT-6 Astra at 41.4% on AutomationBench, ahead of Opus 5.5.

    GPT-6 Sol costs half as much, and that is its whole argument

    At API rates, Opus 5.5 costs 4 dollars per million input tokens and 20 per million output tokens. GPT-6 Sol costs 2 and 10, after OpenAI cut its price by half from GPT-5.6 Sol. OpenAI also publishes a cost per task on AutomationBench: 27 cents for Sol at xhigh, with Opus 5 at 11.1 times that. Anthropic publishes no cost per task for Opus 5.5. The bill also depends on steps, because a model that needs fewer of them spends fewer tokens. That is exactly what Anthropic's early testers report: GitHub says Opus 5.5 solved more terminal tasks than Opus 5 in less than half the steps, and Lovable reports a third to half fewer steps. How much of the price gap that closes on your work is something only a run on your work can show. Subscribers pay a third bill, the usage meter. Anthropic raised five-hour limits on Pro, Max, Team and seat-based Enterprise, and gave subscription users a reset they can save and spend when they choose. OpenAI promises higher usage limits with Sol, which is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu, and not yet in the regular chat.

    Going back is cheap for me because my context never lived in either app

    The expensive part of switching is the first week, when the new model knows nothing about the project: the approach you already rejected and why, the client note that changed the brief, the rule you have given twice, the file it must never touch. I avoid that week because my project history lives in plain files on my own machine. The Memory Engine I built reads sessions from ChatGPT Codex, Claude Code, Claude Cowork, Gemini and Cursor every 15 minutes and files them into one folder, so Opus 5.5 can start from the decisions already in those files, including the ones I made in Codex. That is what makes a fair test possible. I am giving Opus 5.5 in Claude Code and GPT-6 Sol in Codex the same unfinished task from my own work, in the same folder, with the same files. For each one I will record whether the result passes its check without my help, how many corrections it needs before I would ship it, how long it takes and how far the usage meter moves. The model that finishes more accepted work per allowance gets the subscription. The copy block is that test, written so you can run it on your own task this week.

    ألغيت اشتراكي في ⁨Claude⁩، و⁨Opus 5.5⁩ يعيدني لأجرّبه: الأرقام في صفّه، والسعر في صفّ ⁨GPT-6 Sol⁩

    ألغيت اشتراكي في ⁨Claude⁩ ونقلت عملي إلى ⁨ChatGPT⁩ لسبب واحد: نماذج ⁨OpenAI⁩ كانت تعطيني نتائج أفضل في عملي الحقيقي. لا ولاء عندي لأي شركة، أدفع لمن يُنجز. ومن قرابة ⁨2.45⁩ مليار ⁨token⁩ استخدمتها حتى منتصف سبتمبر ⁨2026⁩، مرّ نحو ⁨1.5⁩ مليار عبر ⁨ChatGPT⁩.

    وفي يوم واحد أطلقت ⁨Anthropic⁩ نموذج ⁨Opus 5.5⁩ وأطلقت ⁨OpenAI⁩ نموذج ⁨GPT-6 Sol⁩. لا تقارن أيّ من الشركتين نموذجها بنموذج الأخرى، لكن المنشورين يذكران النتيجة نفسها لنموذجين أقدم من ⁨Claude⁩ على اختبار ⁨AutomationBench⁩، الذي يقيس مهام عمل تتنقّل بين تطبيقات كثيرة، وهذا يكفي لوضع الرقمين جنبًا إلى جنب: ⁨Opus 5.5⁩ يسجّل ⁨40.0%⁩ في منشور ⁨Anthropic⁩، و⁨GPT-6 Sol⁩ يسجّل ⁨33.2%⁩ في منشور ⁨OpenAI⁩.

    الأرقام إذن تميل إلى ⁨Opus 5.5⁩. أما السعر فيميل بوضوح إلى ⁨Sol⁩: في ⁨API⁩ يدفع المطوّر دولارين لكل مليون ⁨token⁩ من المدخلات بدل ⁨4⁩ دولارات مع ⁨Opus 5.5⁩، و⁨10⁩ دولارات للمخرجات بدل ⁨20⁩. أي نصف السعر بالضبط.

    ولا يحسم هذا قراري. كل شركة اختبرت نموذجها في ظروفها هي، ومستويات الجهد مختلفة، وجدول ⁨Anthropic⁩ نفسه يضع ⁨GPT-6 Astra⁩، النموذج الأعلى عند ⁨OpenAI⁩، فوق ⁨Opus 5.5⁩. وما يهمني في النهاية سؤال آخر: كم مهمة من مهامي تنتهي مقبولة من دون أن أتدخل؟

    لذلك سأجرّب ⁨Opus 5.5⁩ قبل أن أقرر أي اشتراك أحتفظ به: أعطيه وأعطي ⁨GPT-6 Sol⁩ المهمة نفسها، وأحسب ما يُنجزه كل منهما مقابل ما يستهلكه من حصة الاشتراك. وسأحكم بالعمل الذي يُسلَّم فعلًا. والانتقال بين الأداتين لا يكلّفني أسبوعًا ضائعًا، لأن تاريخ مشاريعي لا يسكن داخل أي منهما.

    تابعني على ⁨Instagram⁩:

  • https://www.instagram.com/ai.with.mo/
  • التسجيل في الدورة: https://halaqa.app/enrollment?course=start-with-ai

    الخطوات

    خرجتُ من ⁨Claude⁩ لأن نتائج ⁨ChatGPT⁩ في عملي كانت أفضل

    حتى منتصف سبتمبر ⁨2026⁩ استخدمت قرابة ⁨2.45⁩ مليار ⁨token⁩ عبر أنظمة الذكاء الاصطناعي كلها. نحو ⁨1.5⁩ مليار منها مرّ عبر ⁨ChatGPT⁩، وقرابة ⁨200⁩ مليون فقط في ⁨Claude Code⁩. هذه أرقام من لوحات الاستهلاك عندي، وتفصيلها في أين ذهبت ⁨2.45⁩ مليار ⁨token⁩. وأكبر تجربة خضتها هناك كانت يوم عمل امتد ⁨13⁩ ساعة، أنهى فيه ⁨45⁩ ⁨agent⁩ عملًا كنت قدّرته بثلاثة أسابيع. خطّطت له مع ⁨GPT-6 Astra⁩ ونفّذته مع ⁨GPT-5.6 Sol⁩، وتجد تفاصيل تلك التجربة هنا. لم تكن في القرار عاطفة. الاشتراك يذهب إلى الأداة التي تُنهي عملي، وبالمعيار نفسه أعود الآن لأختبر ⁨Claude⁩ بعد صدور ⁨Opus 5.5⁩.

    منشورا الإطلاق يسمحان بمقارنة لم تُجرِها أيّ من الشركتين

    لا تضع ⁨Anthropic⁩ نموذجها أمام ⁨GPT-6 Sol⁩، ولا تضع ⁨OpenAI⁩ نموذجها أمام ⁨Opus 5.5⁩. لكن الجدولين يلتقيان عند ⁨AutomationBench⁩، وهو اختبار لمهام عمل تتنقّل بين تطبيقات كثيرة: كلاهما يعطي ⁨Opus 5⁩ نسبة ⁨26.9%⁩، وكلاهما يعطي ⁨Fable 5.1⁩ نسبة ⁨31.4%⁩. وما دام الأساس واحدًا، يمكن أن تقرأ الجدولين معًا. على هذا الأساس تنشر ⁨Anthropic⁩ لنموذج ⁨Opus 5.5⁩ نسبة ⁨40.0%⁩، وتنشر ⁨OpenAI⁩ لنموذج ⁨GPT-6 Sol⁩ نسبة ⁨33.2%⁩ عند مستوى الجهد ⁨xhigh⁩. وفي ⁨FrontierCode⁩، وهو اختبار يسأل هل تصلح تعديلات الكود للدمج في مشروع حقيقي، يسجّل ⁨Opus 5.5⁩ نسبة ⁨54.4%⁩ مقابل ⁨50.3%⁩ لنموذج ⁨Fable 5.1⁩، بينما تكتفي ⁨OpenAI⁩ بالقول إن ⁨Sol⁩ يعادل ⁨Fable 5.1⁩ عند ⁨xhigh⁩، من غير أن تنشر رقمه. ولهذه القراءة حدود لا أتجاهلها. كل شركة شغّلت نموذجها في بيئتها الخاصة، ومستوى الجهد ليس واحدًا في الحالتين. والأهم أن جدول ⁨Anthropic⁩ نفسه يضع ⁨GPT-6 Astra⁩، أقوى نماذج ⁨OpenAI⁩، عند ⁨41.4%⁩، أي فوق ⁨Opus 5.5⁩.

    ⁨GPT-6 Sol⁩ بنصف السعر، لكن عدد الخطوات يدخل في الفاتورة أيضًا

    في ⁨API⁩ يكلّف كل مليون ⁨token⁩ من المدخلات مع ⁨Opus 5.5⁩ أربعة دولارات، وكل مليون من المخرجات ⁨20⁩ دولارًا. أما ⁨GPT-6 Sol⁩ فبدولارين و⁨10⁩ دولارات، بعد أن خفّضت ⁨OpenAI⁩ سعره ⁨50%⁩ عن ⁨GPT-5.6 Sol⁩. وتنشر ⁨OpenAI⁩ كلفة المهمة الواحدة على ⁨AutomationBench⁩: ⁨0.27⁩ دولار مع ⁨Sol⁩ عند ⁨xhigh⁩، و⁨11.1⁩ ضعف ذلك مع ⁨Opus 5⁩. ولا تنشر ⁨Anthropic⁩ رقمًا مقابلًا لنموذج ⁨Opus 5.5⁩. لكن سعر الـ ⁨token⁩ لا يصنع الفاتورة وحده، فالنموذج الذي يصل إلى الحل بخطوات أقل يقرأ ويكتب أقل. تنقل ⁨Anthropic⁩ عن مختبريها الأوائل أن ⁨GitHub⁩ رأت ⁨Opus 5.5⁩ يحل مهام ⁨terminal⁩ أكثر مما يحله ⁨Opus 5⁩، في أقل من نصف الخطوات، وأن ⁨Lovable⁩ رأت الخطوات تقل بما بين الثلث والنصف. هذه المقارنة مع ⁨Opus 5⁩ وليست مع ⁨Sol⁩، فلا أعرف بعد إن كانت تكفي لسد فرق السعر. وأغلب من يقرأ هذا لا يدفع عبر ⁨API⁩ بل باشتراك، والعدّاد هناك مختلف. رفعت ⁨Anthropic⁩ حدود الاستخدام المحسوبة على كل خمس ساعات لباقات ⁨Pro⁩، ⁨Max⁩، ⁨Team⁩ و⁨Enterprise⁩ بنظام المقاعد، ومنحت المشتركين ⁨reset⁩ لحدود الاستخدام يحتفظون به ويستعملونه متى شاؤوا. وتعد ⁨OpenAI⁩ بحدود أعلى مع ⁨Sol⁩، وهو متاح في ⁨ChatGPT Work⁩ و⁨Codex⁩ لمشتركي ⁨Plus⁩، ⁨Pro⁩، ⁨Business⁩، ⁨Enterprise⁩ و⁨Edu⁩، ولم يصل بعد إلى المحادثة العادية.

    يبقى الاشتراك للنموذج الذي يُنهي عملًا مقبولًا أكثر بالحصة نفسها

    الاختبار الذي سأجريه بسيط. آخذ مهمة حقيقية لم تكتمل بعد، وأسلّمها إلى ⁨Opus 5.5⁩ داخل ⁨Claude Code⁩ وإلى ⁨GPT-6 Sol⁩ داخل ⁨Codex⁩، في المجلد نفسه ومع الملفات نفسها. ثم أسجّل لكل منهما ما اجتاز الفحص من عمله دون مساعدة، وعدد التصحيحات التي احتاجتها النتيجة، والوقت الذي استغرقه، والقدر الذي استهلكه من عدّاد الاشتراك. النموذج الذي يُنهي عملًا مقبولًا أكثر بالحصة نفسها يحتفظ بالاشتراك. التعليمات جاهزة في مربع النسخ أسفل الصفحة، ويمكنك تشغيلها على مهمة من مهامك. أغلى ما في تغيير الأداة هو الأسبوع الأول، حين لا يعرف النموذج الجديد شيئًا عن مشروعك: الحلول التي جرّبتها ورفضتها ولماذا، وملاحظة العميل التي غيّرت المطلوب، والقواعد التي كرّرتها على النموذج السابق مرة بعد مرة، والملفات التي طلبت منه ألا يلمسها. أنا لا أدفع هذه الكلفة، لأن تاريخ مشاريعي محفوظ في ملفات نصية عادية على جهازي. بنيت محرّك الذاكرة لهذا الغرض: كل ⁨15⁩ دقيقة يقرأ الجلسات من ⁨ChatGPT Codex⁩، ⁨Claude Code⁩، ⁨Claude Cowork⁩، ⁨Gemini⁩ و⁨Cursor⁩، ويحفظها نصًا عاديًا في مجلد واحد على الجهاز، فيقرؤه أيّ منها بعد ذلك. فحين أفتح ⁨Claude Code⁩ للاختبار، يجد ⁨Opus 5.5⁩ تاريخ المهمة أمامه من اليوم الأول.

    Prompt

    # OPUS 5.5 VS GPT-6 SOL: THE ONE-TASK TEST
    # Run this before you renew, upgrade or cancel either subscription.
    
    # --- PICK THE TASK ---
    # One real, unfinished task from your own work, with a check you trust:
    # a test suite, or a brief a client would sign off.
    # Pick something you would otherwise do yourself this week.
    
    # --- GIVE BOTH THE SAME START ---
    # Put the context in files both tools read:
    # AGENTS.md for your standing rules and how finished work is checked.
    # STATE.md for where the work is, what was decided and why, what was
    # rejected, and what comes next.
    # Codex reads AGENTS.md by default. Claude Code reads it when the folder has
    # no CLAUDE.md; otherwise put "@AGENTS.md" on the first line of CLAUDE.md.
    #
    # Open Claude Code on Opus 5.5 and Codex on GPT-6 Sol, each on its own branch
    # or copy of the folder. Send both the same first message:
    #
    # Read AGENTS.md, then STATE.md. Finish the task under "Next" in STATE.md.
    # Run the checks named there before you tell me you are done. When you finish,
    # rewrite STATE.md so another assistant could continue without asking me.
    
    # --- RECORD FOR EACH ---
    # 1. Did it pass the check without my help?
    # 2. How many corrections before I would ship it?
    # 3. Wall time from the first message to an accepted result.
    # 4. The usage meter before and after the run.
    
    # --- DECIDE ---
    # Keep the tool that finishes more accepted work per allowance on the task
    # you do most often. A close result is a tie: run a second task of a
    # different kind. Run the test again when either company changes a model
    # or a limit.