I wanted to extract some crime statistics broken by the type of crime and different populations, all of course normalized by the population size. I got a nice set of tables summarizing the data for each year that I requested.

When I shared these summaries I was told this is entirely unreliable due to hallucinations. So my question to you is how common of a problem this is?

I compared results from Chat GPT-4, Copilot and Grok and the results are the same (Gemini says the data is unavailable, btw :)

So is are LLMs reliable for research like that?

  • rickdg
    link
    fedilink
    arrow-up
    10
    arrow-down
    1
    ·
    4 months ago

    Treat it like an eager impressionable intern with a confident stride.

      • rickdg
        link
        fedilink
        arrow-up
        2
        ·
        4 months ago

        Proof that you can do anything if somebody piles billions of money on you.