Wiki / concepts / wiki

Verify like a junior analyst

andrej-karpathy's stance on trusting LLM output: treat it as the work of "a very, very junior data analyst… a little bit absent-minded and not quite right all the time." It's useful and fast, but it's a first draft.

His demos include the failures deliberately:

  • ChatGPT Advanced Data Analysis quietly filled a missing 2015 valuation with $100M and reported a 2030 extrapolation of ~$1.7T while the variable actually held ~$20T. Asked about it, the model said "sorry, I messed up."
  • Deep research's table of US LLM labs left out xAI and included Hugging Face and EleutherAI.
  • Models without a code tool give near-miss answers to large multiplications.

Habits: read the generated code, click the citations, transcribe an image to text first so you can check what the model saw before asking questions, and weight trust by how common the knowledge is (blood-test ranges are fine as a first draft; still see a doctor). It extends his Deep Dive closing line, "check their work, and own the product of your work", and matches the human-owns-verification conclusion of matt-pocock and dhh.

Source: report

Linked from

Andrej KarpathyTool selection by freshness