Making AI Build a Long Run Progression That Holds Up
Ask ChatGPT for a 16-week marathon plan and you will get one in about forty seconds. It will look right. Long runs on Sunday, midweek tempo, easy days between, a taper that starts three weeks out. Open the long run column on its own, though, and the thing tends to fall apart: 10 miles, 12, 14, 16, 18, 20, 22, week after week, no step-backs, no relationship to what the rest of the week is doing.
That column is where most self-built plans break. Not the tempo pace, not the strides, not whether the cutback week lands on week 4 or week 5. The long run is the single biggest stress in a recreational runner’s week, and AI models default to arithmetic progression because arithmetic progression is what the training plans in their training data look like at a glance.
The fix is not a better model. It’s constraints. Write three of them into your prompt and the plan comes back defensible.
The three constraints, stated precisely
1. Cap the long run at 30% of weekly volume. 35% absolute ceiling during peak weeks.
This is the one that catches almost everything else. A 22-mile long run inside a 45-mile week is 49% of your volume in a single session. Your legs experience that as a race, and you recover from it like a race, which means the following week is compromised before it starts.
The same 22-miler inside a 75-mile week is 29%. Same run, completely different training stimulus, because the surrounding fitness absorbs it.
2. Step back every third week by 20-25% on the long run, and drop weekly volume with it.
Not every fourth week. Three weeks of loading is about as long as a runner training around a job can absorb before accumulated fatigue starts outrunning adaptation, and 3:1 is a legacy of elite blocks where the athlete naps in the afternoon. And crucially: the step-back has to apply to the week, not just the long run. Cutting Sunday from 16 to 12 while leaving Tuesday’s intervals and Thursday’s tempo untouched is not a recovery week, it’s a slightly easier Sunday.
3. No quality inside the long run until the final third of the block.
Marathon-pace segments, progression finishes, fast-finish long runs: all of these work, and all of them are the last thing you add, not the first. In a 16-week block, weeks 1 through 10 are aerobic. From week 11, the long run starts carrying race-specific work, and by then the distance stops growing. You build the container first, then you put something in it.
What an unconstrained plan actually returns
Here is the long run column from a 16-week marathon plan generated with a plain prompt: “Build me a 16-week marathon plan, I’m currently running about 30 miles a week and want to run 3:45.” No constraints given.
Week Long run Weekly LR as % Notes from plan
1 10 32 31% Easy
2 12 34 35% Easy
3 14 36 39% Easy
4 10 30 33% Recovery week
5 15 38 39% Easy
6 16 40 40% Last 3 at MP
7 18 42 43% Easy
8 12 34 35% Recovery week
9 18 44 41% Last 5 at MP
10 20 46 43% Easy
11 20 46 43% Last 6 at MP
12 14 38 37% Recovery week
13 22 48 46% Easy
14 20 46 43% Last 8 at MP
15 16 40 40% Taper
16 26.2 35 - Race
Test it against the three constraints and the failures are not subtle.
Every single loading week breaches the 30% cap. Week 13 hits 46%. Nine weeks sit at 40% or above. The plan is not a marathon plan with a supporting week around it, it’s a weekly long run with some filler.
Step-backs arrive on a 4:1 cycle (weeks 4, 8, 12), and the cuts are inconsistent: week 4 drops 29%, week 8 drops 33%, week 12 drops 30%. That’s closer to a crash than a step-back, and it leaves three consecutive loading weeks each time where the long run climbs 2 miles a week with no absorption.
Quality appears in week 6. That is week 6 of 16, with 3 miles at marathon pace tacked onto a 16-miler that is already 40% of the week. By week 14, this plan asks for 8 miles at marathon pace at the end of a 20-miler inside a 46-mile week. Elite blocks contain that session. They contain it at 100+ miles a week.
And the peak: 22 miles in week 13, three weeks out, on 48 miles a week. For a 3:45 marathoner running roughly 10:30 easy pace, that’s a session pushing 3 hours 50 minutes. Longer than the race. On tired legs, four weeks out from the thing they’ve paid for.
The prompt that fixes it
Same goal, same starting fitness, constraints written in:
Build a 16-week marathon plan. Current volume 30 miles/week over 4 runs, goal 3:45 (8:35/mile marathon pace, easy pace 10:15-10:45).
Hard constraints on the long run:
- Long run never exceeds 30% of that week’s total mileage, 35% absolute ceiling in the two peak weeks only
- Weekly volume increases no more than 10% week over week
- Step back every third week: reduce long run 20-25% AND reduce weekly volume 15-20%
- No marathon-pace or progression work in the long run before week 11
- Longest run capped at 20 miles or 3 hours, whichever comes first
- Peak long run lands no later than week 13, then taper
Output as a table with columns: week, long run distance, total weekly mileage, long run as a percentage of weekly mileage, and the session structure for the long run. Flag any week where the percentage exceeds 30%.
That last line is the one people skip. Asking the model to compute and display the ratio forces it to check its own work, and it will often correct a breach mid-table rather than hand you the number. The constraint you can see is the constraint that gets held.
What comes back:
Week Long run Weekly LR as % Long run structure
1 9 33 27% Easy
2 10 36 28% Easy
3 11 39 28% Easy
4 8 31 26% Easy (step back)
5 11 40 28% Easy
6 12 43 28% Easy
7 13 46 28% Easy
8 10 37 27% Easy (step back)
9 14 48 29% Easy
10 15 51 29% Easy
11 16 54 30% 10 easy + 6 @ 8:35
12 12 43 28% Easy (step back)
13 18 52 35% 8 easy + 8 @ 8:35 + 2 easy
14 16 50 32% 12 easy + 4 @ 8:35
15 12 38 32% Easy
16 26.2 34 - Race
Two things to notice. The long runs are shorter and the weeks are bigger. Peak weekly volume went from 48 to 54, while the peak long run came down from 22 to 18. The total training load went up; the single-session risk went down. That trade is the entire point.
The other: quality is late and contained. The first marathon-pace work appears in week 11, by which point the athlete has ten weeks of aerobic base behind it. The hardest session in the plan (week 13’s 8 miles at pace inside an 18-miler) sits in a 52-mile week, which is 35%, right at the declared ceiling and flagged as such.
Checking a plan you already have
Most people reading this are not starting from scratch. They’re partway through something Runna or Garmin Coach or an earlier ChatGPT session handed them, and they want to know whether to trust it.
You don’t need to rebuild anything. Pull the plan into a spreadsheet, or paste it back into the model, and get four numbers per week: long run distance, weekly total, the ratio, and whether the week is a loading or step-back week. Then look for three specific failures.
First, consecutive weeks above 33%. One week is a peak session. Four in a row is a plan that has quietly become long-run-and-recovery, with the midweek runs reduced to jogging between Sundays.
Second, step-backs that only cut the long run. Check the weekly total on your step-back weeks. In the unconstrained table above, week 4 does drop weekly volume from 36 to 30, which is fine, but a lot of app-generated plans hold weekly mileage flat and just shorten Sunday. Your Garmin’s acute training load will show you this straight away: if the 7-day load on a “recovery” week is within 10% of the loading week before it, you did not recover.
Third, quality creeping forward. If marathon-pace segments show up in the long run before you’re two-thirds through the block, the plan is spending race-specific fitness it hasn’t built yet. This is the most common failure in AI plans built for a stated goal time, because naming a goal pace in the prompt makes the model want to use it everywhere.
Strava’s fitness and freshness chart is genuinely useful here, though not in the way most people use it. Don’t watch the Fitness line. Watch Form: if it sits below -30 for more than about ten days at a stretch across a build, something in the loading pattern is not letting you back up.
The bit AI consistently gets wrong even with good constraints
Time, not distance.
A 3:45 marathoner and a 4:45 marathoner following the same 18-mile long run are doing wildly different sessions: 3:05 versus about 4:10. The second one is well past the point where a long run is producing useful aerobic adaptation and firmly into the territory where it’s producing damage and a wrecked Monday.
Models default to distance because plans are published in distance. So state the cap in both: “longest run capped at 20 miles or 3 hours, whichever comes first.” For anyone targeting slower than about 4:15, the time cap will bind first, and the plan should show it. If you ask for a 20-miler and your honest easy pace is 12:00/mile, you’re asking for a four-hour training session, and no amount of good structuring elsewhere saves that week.
The same applies to the half and the 5K. A 5K plan doesn’t need a 30% cap argued from scratch, but the logic holds: the long run in a 5K block is aerobic support, not the main event, and if it’s eating 35% of a 25-mile week you have a plan that is slowly turning into a half-marathon plan. For a broader treatment of how to structure the whole prompt (paces, session types, how to give the model your actual training history rather than a vague fitness description), the guide to building your own plan with ChatGPT and Claude covers the framework these long-run constraints slot into.
One last thing worth writing into the prompt: ask for the reasoning. “For each step-back week, state in one line why the reduction is the size it is.” Models that have to justify a number pick better numbers, and you get something you can argue with on a Tuesday evening when your legs say one thing and the spreadsheet says another. The spreadsheet does not know you slept badly for four nights. You do, and a plan whose logic is visible is a plan you can adjust without throwing out.