AI Running

Five Things AI Gets Wrong When It Writes a Running Plan

Ask ChatGPT for a 16-week marathon plan and you will get one. It will arrive in about forty seconds, formatted into a neat table, with a confident little note at the bottom about listening to your body. It will look exactly like a training plan. That is the problem.

The failure modes are not random. After reading a few hundred of these, generated by GPT-5, Claude, Gemini and the plan engines inside Runna and Garmin Coach, the same five defects turn up again and again. They are structural, they are predictable, and every one of them is visible in the plan text before you run a single step of it. Here is what to look for.

1. Mileage that jumps like a staircase, not a ramp

The classic. You tell the model you currently run 20 miles a week and want to do a marathon in 16 weeks. It knows the peak should be somewhere near 40-45 miles. It has 16 rows to fill. So it fills them.

Here is a real week-by-week total from a ChatGPT-generated plan a runner sent me in August:

Wk1  22    Wk5  34    Wk9   44    Wk13  48
Wk2  26    Wk6  38    Wk10  46    Wk14  38
Wk3  30    Wk7  40    Wk11  48    Wk15  28
Wk4  32    Wk8  42    Wk12  50    Wk16  26 (race)

Week 1 to week 2 is an 18% jump. Week 2 to week 3 is 15%. The runner gained 12 miles in three weeks, which is a 55% increase in load in under a month. The old 10% rule is cruder than the research supports, but the direction is right: acute jumps of this size are where tibial stress reactions and Achilles tendinopathy come from, and they come at week 5 or 6, not week 2, which is why people think the plan was working.

The check: put the weekly totals in a column and compute the week-over-week percentage. Anything over 12% two weeks running is a flag. Anything over 15% in a single week when you are already above 30 miles is a bigger one. Garmin Coach is actually decent at this because it’s building from your real running history; a ChatGPT plan has no history to build from, so it builds from the destination backwards.

Watch also for the first week being too high. Models routinely open a plan at or above your current volume, which means week 1 is already an increase before any progression starts.

2. Three quality sessions, every week, forever

This one is almost universal in LLM-generated plans, and it comes from a specific place: the training literature the models absorbed is heavy on workout descriptions and light on the boring runs between them. Tempo runs, intervals and long runs with pace changes are what people write articles about. Nobody writes an article about Tuesday’s 5 easy miles.

So you get weeks like this, which came out of a “Claude, build me a half marathon plan” session:

DaySession
MonRest
Tue6 × 800m @ 5K pace, 90s rec
Wed5 miles easy
Thu4 mile tempo @ half pace
FriRest
Sat4 miles easy + 6 × 100m strides
Sun12 miles with last 4 @ marathon pace

Count the hard days. Tuesday, Thursday, Sunday: three. The strides on Saturday are a fourth stimulus, mild but not nothing. That leaves two genuinely easy runs in a seven-day week, against three sessions that each need 48 hours of recovery. For a recreational runner on 35 miles a week, this is a plan for getting progressively more tired for five weeks and then getting injured or sick.

The ratio you want is roughly 80/20: about four-fifths of your weekly minutes at a conversational effort, one-fifth hard. Two quality days per week is the standard for a reason, and plenty of good marathon plans run on one quality day plus a long run.

The check: for each week, count sessions that are not “easy” or “rest”. If the answer is 3 or more for more than a couple of weeks in the block, the plan is over-cooked. Then check your own Strava: if your “easy” runs are averaging within 45 seconds per mile of your tempo pace, you have a four-hard-day week regardless of what the plan calls them.

3. No down weeks, or down weeks that aren’t down

Look back at that mileage staircase in section 1. Weeks 1 through 13 go up, up, up, with no recovery week anywhere. Fourteen onward is the taper. That’s thirteen consecutive weeks of increasing load, which no coach on earth would write.

Adaptation happens during recovery, not during the work. A structured block wants a step-back every third or fourth week: volume drops 20-30%, intensity drops with it, and you arrive at the next block absorbing the previous one rather than accumulating fatigue on top of it.

The AI failure here has two flavours. The first is the plan above: no down week at all, because the model was filling a grid from A to B and a dip looks like a mistake in a grid. The second is subtler and more common in the better outputs. The plan includes a week labelled “recovery week” or “cutback”, the mileage drops from 44 to 40 (a 9% reduction, which is noise), and the workouts stay identical. That is not a down week. That is a normal week with a comforting label.

Runna handles this properly most of the time. ChatGPT plans need checking every time.

The check: find every week the plan calls recovery, cutback, or down. Compute the drop from the week before. Under 20%, it isn’t one. Then check the workouts: if the recovery week still contains a threshold session and a long run at 90% of the previous week’s distance, it isn’t one either. And if you count more than four consecutive weeks of rising volume anywhere in the plan, ask for the block to be restructured.

4. Paces invented out of nothing

This is the one that bothers me most, because it looks the most authoritative. You get a table like:

Easy:        9:15–9:45 /mi
Marathon:    8:20 /mi
Threshold:   7:45 /mi
Interval:    7:10 /mi
Repetition:  6:45 /mi

Beautifully structured. Correctly ordered. Plausible gaps between the zones. And, if you never told the model a recent race result or a tested threshold, entirely fabricated. The model has a statistical sense of what a set of training paces looks like, and it generated one that looks right. It has no idea whether you can hold 7:45.

Test it against a real equivalence. If you ran a 10K in 48:00 (7:44/mi), your threshold pace is roughly 7:50-8:00, your marathon pace is nearer 8:45-8:55, and an easy pace of 9:15 is too fast, not too slow. The table above belongs to a runner about four minutes quicker over 10K than the one it was handed to.

The downstream damage is specific: you run your easy days at marathon effort, your marathon-pace long runs at threshold, and your intervals in a zone that produces a lot of fatigue and not much adaptation. Then the plan gets blamed for the injury when the paces were the problem.

Worse: models do this silently even when you do give them data, because they convert badly. Ask GPT-5 to derive marathon pace from a 21:30 5K and you will often get something 15-20 seconds per mile optimistic, because the raw equivalence tables assume marathon-specific endurance you may not have on 30 miles a week. Race-equivalent is a ceiling, not a target.

The check: take any pace the plan gives you, find your most recent hard race or time trial, and run it through a real calculator (VDOT, Riegel, whatever you trust) yourself. If the plan’s paces are more than about 10 seconds per mile faster than the calculator’s, the model guessed. If you never gave it a race result, it definitely guessed, and you should go and run a parkrun before using the plan at all. Our guide on building your own plan with ChatGPT and Claude walks through exactly which inputs to feed the model so the paces come out anchored to your actual fitness.

5. A taper that starts a week too late and cuts the wrong thing

Tapering is where AI plans get squeezed by their own arithmetic. The model has allocated 16 weeks, it wants a big peak week to look impressive, and the taper is whatever is left over. So the 20-miler lands on week 14 of 16, and you get ten days of frantic reduction before the race.

Two distinct errors show up. The first is length: a marathon taper wants two and a half to three weeks, a half wants ten to fourteen days, a 5K wants about a week. AI plans routinely give a marathon a 10-day taper, often with the peak long run only 13 days out. You arrive at the start line with your legs still processing week 14.

The second is what gets cut. A good taper drops volume hard (down to roughly 50-60% of peak in the final week) while keeping intensity: a few race-pace miles, some strides, enough sharpness to keep the legs awake. AI plans often do the reverse. They cut the workouts entirely and keep a fair amount of easy mileage, so you show up rested and flat, with legs that have not run fast in seventeen days.

Here is the difference on the final ten days before a marathon:

DayTypical AI taperWhat you want
-108 easy8 easy
-96 easy5 easy + 4 × 90s @ MP
-8Rest5 easy
-710 easy8 with 3 @ MP
-65 easy4 easy
-5RestRest
-46 easy5 easy + 6 strides
-34 easy3 easy
-2RestRest
-13 easy2 easy + 3 strides

Same rough volume. Completely different legs on Sunday.

The check: count backwards from race day to the final long run. For a marathon it should be 21 days out at minimum, and the final three weeks should read roughly 75%, 55%, 35% of peak volume including the race. Then check that at least two sessions in the last fourteen days contain something at goal pace or faster. If the last quality work you do is 16 days before the start, ask for it back.

Running the audit

None of this takes long. Paste your generated plan into a spreadsheet, one row per week, four columns: total mileage, week-over-week change, count of quality days, and whether it’s a down week. Five minutes of arithmetic surfaces four of the five problems above. The fifth, the paces, needs a race result and a calculator.

And when you find a problem, go back to the model and say so specifically. “This plan has no recovery weeks; restructure it as 3 weeks up, 1 week down, keeping the same peak” produces a good revision almost every time. The models are genuinely capable of building sound training structure. They just don’t do it unprompted, because nothing in how they generate a plan forces them to check the plan against itself.

That last part is worth sitting with. The plan is not wrong because the model is stupid. It’s wrong because it was writing a document that looks like a training plan, and you were reading it as a set of instructions for your tendons.