Would an AI lie to you?


I don’t recall ever seeing services that offered to do all your research for you, and to write the paper for you, so now I’m feeling old. Such services exist, though. One is called Research Gold — you just tell them your thesis, they dig into the library, and they return a complete, fully written paper that you can then submit to win academic honor and glory.

Research Gold, a site that advertises services for advertises services for medical researchers, including drafting peer-review ready manuscripts , including systemic reviews and meta-analyses, claims that it’s “100% human-written, never AI,” and lists a number of PhD reviewers and professional methodologists on staff that carry out this meticulous, difficult work.

There’s something else I haven’t seen before, that peculiar need to reassure customers that nosir, I’m no AI. I am a human. You can trust me. Except…

The problem: The PhD reviewers Research Gold lists on its site are AI-generated and don’t exist. Other methodologists it lists are real, but are not aware their identity is being used by Research Gold. When I tried calling the company, an AI agent that refused to concede it was AI answered and kept trying to sell me Research Gold’s services. Email and chat communication with the company were also AI generated.

It’s 100% AI through and through! But I might be tempted to use an AI to brute-force its way through a literature review — I’ve done the thing where you get on PubMed and throw keywords at it to try and find all the papers relevant to your interests. It’s the kind of novel search I’ve thought could be automated, and which would churn out a long list of potential papers for human review and evaluation. I’m not the only one who has thought that.

Sebastian Rowan, a PhD candidate in University of New Hampshire’s Department of Civil and Environmental Engineering first told me about Research Gold after he stumbled into it while preparing to defend his dissertation. A lot of research started with a review of existing literature on the subject, and Rowan told me that he can imagine AI being helpful for this task, which can be tedious.

The problem is hallucinations — an AI will blithely make up citations to papers that don’t exist. Nowadays, when I review papers, I have to check all the citations to see if they were real, which never used to be a problem.

So even at a task where you would think an AI would excel, it makes the problem even worse.

No thanks, Research Gold. Change your name to Research Shit.

Comments

  1. Walter Solomon says

    So, will an AI-written research paper cover the shortcomings of AI such as those hallucinations?

    And if it doesn’t, would that indicate the AI is being intentionally misleading? If it’s being intentionally misleading, does that mean it’s self-aware?

  2. says

    Can a parrot lie if it doesn’t understand what it’s saying?

    Meanwhile, today’s an art practice day, as soon as I can get Glaze and Nightshade downloaded.

  3. anxionnat says

    Such services (pay me to write your paper/thesis for you, to take your exams for you) did exist, even as early as the 1980s, depending on the discipline. (I know, because I wrote a master’s thesis for a person I shall not name, though my nearest and dearest know who it was. I also wrote all the documentation/field work that supported the thesis, over a period of about two years.) Most of this work was sort of sub rosa, as ’twere. No AI involved. There will always be people who cheat, who are willing and able to pay others to do real work for them. And, yeah, there are people like me who have non-marketable reading and research skills, who will take such jobs. Pays (a bit) better than flipping burgers at Mcdonalds.

  4. Bruce says

    My guess is that in the future, when a journal editor receives a submitted manuscript, the first thing they will do is to submit it to a service that automatically uses several AI engines to look up all the references, and each AI engine separately reports how many papers it could verify.
    If I were the journal editor, and every AI engine found almost all the references, then I’d email the author of the submitted manuscript and ask for copies of the few referenced articles I couldn’t easily verify. I wouldn’t submit the article to even start peer review until I got back all the papers I hadn’t found. And if a manuscript had several references that were not confirmed, I wouldn’t even bother doing that. And I would have my assistant or secretary use AI to automate all of the above, as much as possible.
    So a paper with a few typos in the references could save itself, but a paper by AI would be rejected without even peer review.

  5. raven says

    Would an AI lie to you?

    Yeah, they do it all the time.
    More specifically, they will hallucinate and give you an answer, that may or may not have anything to do with reality.

    I use Google Search a lot and the AI frequently answers at the top of the search results.
    I always check the AI to make sure it is probably right. Sometimes it is wildly wrong.
    The AI is biased to give you and answer so if it can’t find the answer, it will make something up instead.

    I asked once for a news story about an old rock music singer who was arrested for indecent behavior at a concert once. She beat the charge in court. This was more an example of clueless law enforcement anyway.

    The AI kept giving me fake court cases because I couldn’t quite remember the name and kept giving it the name of another singer.
    Finally, I guessed right and got the answer.

    Wendy O. Williams.

    Wendy O. Williams was not arrested in Cleveland, but rather charged and put on trial following a January 21, 1981, concert at the Agora Ballroom. She was charged with pandering obscenity for performing wearing only shaving cream and black tape over her nipples, and was acquitted by a jury in April 1981.

    I remember this because I thought it was really stupid when it came out in the newspapers in 1981. A violation of the First Amendment among other things.

    She did not have a happy ending, dying from suicide in 1998.

  6. robro says

    AI might be helpful, but I hope no smart person is going to rely on and submit a thesis written entirely by an “AI”, i.e. some form of Large Language Model. The results aren’t trustworthy. I assume the people evaluating the results can spot the signs of an LLM written thesis without too much difficulty. Even as a research tool LLMs aren’t 100% reliable…you need to validate all sources and citations because faking citations seems to be one of the more common problems with them. The LLM aims to please, even if it has to make shit up. Then, so I’ve heard, there’s orals and/or the defense of your thesis which is live. If you’ve relied on LLMs too much you’ll just look like an ignorant idiot and fail that part because there’s no faking it.

  7. Akira MacKenzie says

    The thing is “AI” isn’t AI. As I understand it (and I could be wrong), Large language models composes its answers by stringing together the words that are most likely to appear next. They don’t have the ability to check its work and determine truth from falsehood.

  8. rickarddavidt says

    Would an AI lie to you?
    Technically speaking, no, since lying at least implies intent and knowledge of the truth, neither of which are possessed by the stochastic parrots misnamed AI.
    Now, the companies that own and sell “AI” are another thing.

  9. says

    Wait! So, we should use AI to check if a paper was generated by AI?
    Isn’t that like asking tRUMP to check and tell you if one of his magat minions is lying? A mental circlejerk ensues.
    Therefore, we can only conclude that all the drooling magat cult members are truly ‘artificially Intelligent’.

  10. says

    @11: I’d say it’d depend on the nature of the AI. Remember the term’s become useless precisely because it’s been conflated with generative AI and LLMs, rather than all the other ways we can teach a computer to do stuff.

  11. says

    Curiosity, since I’ve got Nightshade and Glaze: Glaze poisons the data so a computer can’t replicate your art style, and Nightshade poisons the metadata to confuse the AI what it’s a picture of…

    Is there a text-based version that poisons your writing style or the type of prose you’re writing?

  12. Akira MacKenzie says

    The problem isn’t that AI lies, it’s that people assume it’s giving them the truth.

  13. says

    @13 Recursive Rabbit asked about poisoning writing style.
    I reply: As a user of dinosaur computers and obscure operating systems, I have found that converting image files to an older type of graphic file deletes the metadata. And, while I don’t have knowledge of how sneaky AI is, would using an unconventional font or converting a page of text to an image reduce the likelihood of AI being able to use it? Also, there is the possible use of steganography for obscuring things.

  14. says

    One other obfuscation technique would be to edit your images and text and insert a few letters the same color as the background randomly. Think of using a white text letter instead of a spaces between words to jumble things.
    Good luck, it’s hard to prevent intellectual property theft in today’s world. Even copyright registration with the obligatory uploading of the work makes it vulnerable to harvesting theft.
    (I won’t post any more comments here since I don’t want to hijack this thread)

  15. CompulsoryAccount7746, Sky Captain says

    Wikipedia – On Bullshit

    a theory of bullshit […] The liar wants to steer people away from discovering the truth, and the person telling the truth wants to present the truth. Bullshitters differ from both […] Persons who communicate bullshit are not interested in whether what they say is true or false, only in its suitability for their purpose.
    […]
    Some researchers believe that Frankfurt’s concept suitably describes the behavior of large language model-based chatbots and is more accurate than terms like “hallucination” or “confabulation”. The uncritical use of LLM output is sometimes called botshit.

  16. asclepias says

    Spare me from AI! Every time I try to read a paper online, Gemini comes up with a ‘let’s get started! Would you like me to summarize this paper for you?’ No, no let’s not. First of all, there’s no we, there’s just me, and if I don’t read the paper myself, I haven’t learned anything. What’s the point of that? And as for summaries, the authors already did that. It’s called the abstract! all swear words running around in my brain have been left out of this reply

  17. John Morales says

    the stochastic parrots misnamed AI

    Heh. That idea is still around, I see.

    Anyway.
    That AI is being used by scammers does not entail that AI is a scam.

  18. Elladan says

    I think saying the AI is hallucinating when it’s making shit up is giving it way too much credit, and isn’t accurate.

    The AI isn’t hallucinating when it’s wrong. It’s hallucinating when it’s doing anything at all. Hallucinating is all it can do. It’s just that through advances in computation, the hallucinations are almost always grammatically plausible and occasionally contain correct statements too.

    They aren’t thinking clearly and then occasionally hallucinating. They’re hallucinating all the time! That’s all they do!

  19. John Morales says

    Elladan, you are taking a term relating to a specific failure mode of LLMs (which are not AI in general) and claiming it’s all it ever does, which is contrary to the technical usage. I assure you, they are not always failing thus.

    (Are you imagining it’s the same thing as ‘hallucination’ for humans?)

    LLMs have no volition or intent, so they cannot ‘intend’ to lie.
    But they can fabricate stuff and ignore stuff, to the effect that their output is false.

    FWIW, here is a list of ‘definitions’ in use, all of which I verified (got the URLs):
    Ji et al., 2023 — “Hallucination refers to generated content that is unfaithful to the source or unsupported by evidence.”
    Maynez et al., 2020 — “Hallucinations are outputs that contain information not present in the input.”
    Bang et al., 2023 — “A hallucination is an incorrect or fabricated statement produced with unwarranted confidence.”
    Rawte et al., 2023 — “Hallucination is the generation of content that is factually incorrect or nonsensical despite appearing plausible.”
    Min et al., 2023 — “Hallucination is the production of statements not supported by training data or external evidence.”
    Huang et al., 2023 — “Hallucination denotes outputs that deviate from the ground truth or input context.”
    Krishna et al., 2023 — “Hallucinations are confident assertions of falsehoods.”

  20. robro says

    Akira MacKenzie @ #7 — As I understand it, LLMs don’t operate on words. They are a “next token predictor”. They process tokens, which might be a word or part of a word or, I suppose, several words.

  21. beholder says

    Would an anti-AI fanatic lie to me?

    Judging from all the overconfident bullshit in this comment section, most of you would do well to open up an LLM chat bot and test your hypotheses directly.

    So even at a task where you would think an AI would excel, it makes the problem even worse. — @PZ

    Yes, it’s possible to use a good tool badly. That’s not an indictment of the tool.

    Can a parrot lie if it doesn’t understand what it’s saying? — @2 RR Rabbit

    Technically speaking, no, since lying at least implies intent and knowledge of the truth, neither of which are possessed by the stochastic parrots misnamed AI. — @9 rickarddavidt

    Your knowledge on the subject is hopelessly outdated and bears no resemblance to the way modern artificial neural networks run when they’re using the transformer model.

    They don’t have the ability to check its work — @7 Akira

    So use another LLM to interrogate its intermediate output stages. This is not a novel approach.

    Wait! So, we should use AI to check if a paper was generated by AI? — @11 shermanj

    Well, yes. It’s a lot faster than checking with a human.

    Is there a text-based version that poisons your writing style or the type of prose you’re writing? — @13 RR Rabbit

    It’s hilarious that you think a poisoned writing style would be anything a human would want to read. Be careful what you wish for.

  22. John Morales says

    cf. https://aclanthology.org/2023.emnlp-main.614.pdf

    Tokenization in LMs
    Tokenization — segmenting text into atomic units — is an active research area. Proposed approaches range from defining tokens as whitespace‑delimited words (for languages that use whitespace), which makes the vocabulary extremely large, to defining tokens as characters or bytes, making the tokenized sequences extremely long in terms of number of tokens; see Mielke et al. (2021) for a detailed survey. A commonly used solution now is to tokenize text into subword chunks. With Sennrich et al. (2016), one starts with a base vocabulary of only characters, adding new vocabulary items by recursively merging existing ones based on their frequency statistics in the data. Other approaches judge subword candidates to be included in the vocabulary using an LM (Kudo, 2018; Song et al., 2021). For multilingual models containing data in a variety of scripts, even the base vocabulary of only characters (based on Unicode symbols) can be very large with over 130K types. Radford et al. (2019) instead proposed using a byte‑level base vocabulary with only 256 tokens. Termed byte‑level byte pair encoding (BBPE), this approach has become a de facto standard used in most modern language modeling efforts (Brown et al., 2020; Muennighoff et al., 2022; Scao et al., 2022; Black et al., 2022; Rae et al., 2022; Zhang et al., 2022b). In this work, we investigate the impact this tokenization strategy has on LM API cost disparity as well as downstream task performance (i.e., utility) across different languages.

  23. badland says

    They don’t have the ability to check its work — @7 Akira

    So use another LLM to interrogate its intermediate output stages. This is not a novel approach.

    Use a hallucinating lie machine to evaluate the output of another hallucinating lie machine! beholder your genius knows no bounds.

  24. Fez says

    The institution I used to be associated with had a group that created an LLM-based service to, as I understood it, assist in the production of academic papers – to behave as an organizational assistant rather than a research engine. Not my wheelhouse which is why I ask, but do people think this kind of incorporation of an LLM into a workflow is ok or does it still cross the line because there is still some kind of content generation that didn’t originate in the mind of one of the human researchers?

    (It’s the “Agentic LaTeX” product at https://grailai.io/ if curious; I see they’ve branched out services since I left the org. Including becoming a commercial entity, with all the negatives that might imply)

  25. John Morales says

    Fez, does it improve the workflow? If so, yes, it’s OK, if it worsens it, it is not, and if it makes no diff, who cares?

    FWIW, mathy types are using them for real “work”:

  26. beholder says

    @27 Fez

    Probably one of my more frequent LLM tasks is to generate TeX templates for math equations. I prefer to spend time filling in the formula instead of struggling with the scripting engine as it hiccups over every syntax error I inevitably introduce, were I to try writing TeX off the top of my head.

    As far as I am concerned, LLMs have made TeX accessible.

Leave a Reply