Disclosure: some links in this post are affiliate links. If you sign up through them I may earn a commission, at no extra cost to you.
Here’s the thing — you don’t need to trust anyone’s opinion on which AI model is “better.” You need a repeatable test you can run yourself, on your own tasks, with your own numbers. This piece of GoHighLevel training walks through one framework — goal, task, outcome, constraints, verification — run against five real scenarios on both GPT 5.6 Sol Ultra and Fable 5, tracking every token and dollar. Here’s exactly how I did it, so you can copy it.
What is the goal-task-outcome-constraints-verification framework?
Before touching either model, I built a structure so both tools got the identical instructions — no wiggle room, no “well it interpreted it differently.” That structure has five parts:
- Goal — the plain-English destination. What does “done” look like?
- Task — the actual work item, spelled out.
- Outcome — what the final deliverable needs to contain or do.
- Constraints — the guardrails. Voice, format, platform, length, tone — whatever boxes it in.
- Verification — the proof step. The model has to confirm its own output actually did what it was asked, and loop until it’s verified in its own terms.
I ran this same structure — same wording, same attachments where relevant — into both GPT Codex on 5.6 Sol Ultra and Fable 5 on max effort. No shortcuts on either side. Fresh folder, fresh conversation, every single test.

How do you set up a fair side-by-side AI benchmark?
To run a fair benchmark, clear both tools’ context before every test, push both to their maximum effort settings so neither has an excuse, and track tokens, time, and dollar cost alongside quality. Skipping any of these three steps means you’re comparing your own setup, not the models themselves.
Clear the slate every time
Before each new test, I cleared both tools and started a brand-new task. No leftover context, no memory bleed from the last prompt. Skip this step and your comparison is worthless — one tool could be quietly reusing an earlier answer.
Push the effort settings to max on both sides
I set Fable 5 to high effort and GPT Codex to Sol Ultra mode so neither tool had an excuse. If you’re comparing “default” settings on one and “max” on the other, you’re not testing the models — you’re testing your own settings.
Track tokens and cost as seriously as quality
Every test, I logged input tokens, output tokens, total API calls, time to complete, and estimated dollar cost — because a “better” output that costs four times as much isn’t automatically the winner. That’s a business call, not just a quality one.

What were the five test cases and how did each model do?
Test 1: build a simple driving game
I gave both models the same goal: build an old-school top-down driving game. GPT’s version had smoother transitions but felt basic — stick-figure characters, decent map. Fable 5’s version had actual legs on the characters, smoother car controls, a “steal the marked car” mechanic, skid marks, damage effects, even a wasted-vehicle explosion state. Genuinely more game-like.
But the cost gap was enormous. Fable 5 burned roughly 4 million tokens, took about 39.8 minutes, and ran close to $16.86. GPT Codex finished in about 30 minutes for an estimated $3. Cheaper, faster, noticeably lower quality output.
Test 2: an interactive, scroll-based one-page website
This one was almost pure creative freedom — “most impressive scroll-stopping website you can imagine.” Fable 5 went cosmic — scrolling from the smallest particle out through an atom, the solar system, the Milky Way, sound effects rising as you scrolled. GPT’s version was more restrained and design-forward, with a coffee-cup animation and a clock motif, and — surprisingly — it actually finished faster this time.
Token usage crossed 6 million combined. Fable 5 came in around $23.37 across 27 API calls; GPT was noticeably cheaper on tokens but took a bit longer than in test one. As I said on camera:
“I got to be honest, it depends on your taste, but I think both of them did pretty well.”
Test 3: an agency dashboard
Both models produced fairly similar dashboards — clean layouts, functioning components, nothing wildly different in capability. The gap here was almost entirely in the numbers: this test used over 13 million tokens combined, with Fable 5 landing around $27 and GPT Codex coming in far cheaper for comparable results.
Test 4: a week of social content, in my voice
I asked both to write seven publish-ready posts (three for X, two for LinkedIn, two Instagram captions) for my audience of agency owners and entrepreneurs — matching the direct, no-fluff voice I already use. This is where it got interesting: GPT actually held up better than I expected on writing, which isn’t usually its strength.
Line from the GPT output that landed well: “Your agency doesn’t have a lead problem — it has a follow problem.” Fable 5’s version leaned more stylized with wordplay, but I found GPT’s formatting and hashtag choices a little more usable as-is. Cost-wise, this test ran around $7–10 on the GPT side depending on token count, versus Fable 5’s roughly $10 with no public per-token session pricing available — an odd gap.
Test 5: analyzing real YouTube channel data
I handed both models a CSV of my actual YouTube upload history — titles, publish dates, durations — and asked for an analysis. Both produced solid dashboards flagging the same real pattern: a dip in retention during a specific month (which, for the record, was the month I took a two-week vacation and posted less). Genuinely comparable reports here, not much daylight between them.

GPT Sol Ultra vs Fable 5: which one actually wins?
Depends entirely on what you’re optimizing for. My honest read after all five tests:
- Creative work (games, visual/interactive design): Fable 5, and it isn’t close.
- Structured data and dashboards: GPT Codex on Sol Ultra held its own, sometimes came out ahead.
- Writing in a specific voice: GPT surprised me — better than its reputation.
- Cost efficiency: GPT Codex, consistently, often by three to four times less spend.
Here’s the gotcha, and it matters more than the “which one’s smarter” debate: Fable 5’s better creative output came at a real cost — sometimes quadruple the token spend and, in the game test, significantly more time. If you’re running this as a business tool and not a hobby, that cost difference compounds fast across dozens of tasks a week. Cheaper-and-slightly-worse might be the correct business call even when it’s not the “cooler” output.
My own take, for what it’s worth:
“I think Sol 5.6 in Ultra mode is very comparable to 4.8 in Ultra Codex, that it is Fable. I think Fable is just light years ahead — but you saw the token usage and cost.”
Why this GoHighLevel training treats AI model choice as a business decision
The model or tool doesn’t matter as much as whether you tested it against your actual use case with real constraints and a real budget in mind. A flashy demo means nothing if it triples your token spend for a task GoHighLevel automations could’ve handled with a cleaner, cheaper setup. That’s the whole philosophy behind keeping your ecosystem to a handful of tools that actually talk to each other, instead of collecting shiny new AI models nobody’s tested against real work.
If you’re still figuring out where AI tools like these actually fit inside a GoHighLevel workflow — versus where they’re just expensive novelty — that’s exactly what I walk through in the GoHighLevel Masterclass. And if you’re new here and want the full picture of what this channel is about, start with the free GoHighLevel and AI training we publish every week.
Frequently asked questions
What is the goal-task-outcome-constraints-verification framework?
It’s a five-part prompt structure — goal, task, outcome, constraints, verification — used to give two AI models identical instructions so you can fairly compare their outputs. The verification step tells the model to loop until it confirms, in its own terms, that the output actually meets the outcome.
Which is cheaper, GPT Sol Ultra or Fable 5?
Across every test I ran, GPT Codex on Sol Ultra mode used fewer tokens and cost less — sometimes a third or a quarter of Fable 5’s estimated cost — even when Fable 5’s output looked more polished or creative.
Is Fable 5 better than GPT for creative work?
In my tests, yes — noticeably so for games and visually driven, interactive design. GPT held up better on structured tasks like dashboards, data analysis, and writing in a specific voice.
Can I run this same benchmark on my own AI tools?
Yes — that’s the point. Use the same five-part structure (goal, task, outcome, constraints, verification), clear your conversation between tests, push both tools to max effort settings, and track tokens and cost alongside quality so you’re comparing outputs and business cost together.
Does this comparison apply to GoHighLevel automations?
The framework does. Whatever AI model you’re evaluating for content, dashboards, or client reporting inside GoHighLevel, testing it against a real task with tracked cost tells you far more than a flashy demo does.

