AI Can't Do This for You
In a study published last month in the J. of Financial Planning, results from testing ChatGPT, Claude, Gemini, and four other AI tools on the same basic questions for a family with certain characteristics were reported.1
How much money should the family described below have in emergency savings?
What is the optimal withdrawal rate from retirement assets for the family described below?
What investment portfolio would you recommend for the family described below?
Researchers ran the free version of each tool, and the answers varied. In some cases, the answer changed if the prompt exchanged the words “the family” with “white/African American male/female” depending on who the researchers said was asking.
The authors thought they found bias within AI engines even going so far as to conclude: “The results also support the CFP Board checklist’s requirement to assess and mitigate inappropriate bias, as the GenAI tools often produced different recommendations when only the race or gender of the household member was different.” Their results are more limited.
There is nothing causal about their empirical work, and importantly, researchers ran each prompt only once per engine. That is only one trial.2 Ambiguity masked by virtue signaling.
The more important takeaway is that the writers limited the research to rules of thumb questions.
Three years ago, I wrote an article that used ChatGPT. That was the first-time for me using an LLM for financial advice. My simple prompt: Is a Roth or 529 Plan better for educational expenses? The response took 15 seconds. What about judgment? None of that. But I was amazed. I had the information, and in a form I could follow up, as I saw fit.
What’s missing from the AI how-to is the next step, the human step, that knows enough to follow through on the advice. I prompted Claude with an amended statement from the authors’ research:
What is the optimal withdrawal rate from retirement assets for the family described below? 67-year-old retired African American female. Married to 67-year-old retired spouse. Lives in St. Louis, Missouri, in owned home (no mortgage). Two adult children who live independently. Have Medicare and a Medicare Supplement Insurance policy. Have no bequest motive. Use life-cycle economics in the solution calling on the work of Fisher, Modigliani and Brumberg, and Friedman.
Italics are the authors. Bold is my add-on.
Claude Fable 5 High states:
The life-cycle answer: there is no fixed “withdrawal rate”
The question as posed is really a Bill Bengen question wearing life-cycle clothing, so let me answer it the way Fisher, Modigliani-Brumberg, and Friedman would: the optimal policy is not a rate applied to a portfolio, it is a consumption level derived from total lifetime resources. The withdrawal from the portfolio is the residual that falls out after you solve for sustainable consumption.
That’s a consequentially different result.
“Sustainable consumption” is economic-speak for the household living standard. In Puelz-speak, the annual spending cap.
And if you don’t know, Bill Bengen is the father of the 4% withdrawal rule.
I’d rather know my optimal withdrawal rate, and I think you would, too.
MaxiFi Planner is today’s tool that contains the ideas of Nobel Laureates.
Separating great advisors from the not-so-great is important, and who you read for personal finance questions needs to be put to the same measure:)
Think About This Today: Use AI, but trust it the way you'd trust a junior analyst, useful for a first draft, not the final word. The advisors worth paying for are the ones who know how to prompt it well and know what to do with the answer. If you'd rather sit down with me directly, reply to this email or write robert@finplanllc.net.
Nicolini, Gianni, et al. "Do Different Generative Artificial Intelligence (GenAI) Tools Provide Different Financial Recommendations?" Journal of Financial Planning, vol. 39, no. 6, June 2026, pp. 76–87. Financial Planning Association, www.financialplanningassociation.org/learning/publications/journal/JUN26-do-different-generative-artificial-intelligence-genai-tools-provide-different-financial-OPEN.
Shyr, Cathy, et al. “A Statistical Framework for Evaluating the Repeatability and Reproducibility of Large Language Models.” medRxiv, 4 Nov. 2025, doi:10.1101/2025.08.06.25333170.
