this post was submitted on 04 Dec 2023
698 points (92.8% liked)

Technology

58061 readers
31 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 1 year ago
MODERATORS
 

We demonstrate a situation in which Large Language Models, trained to be helpful, harmless, and honest, can display misaligned behavior and strategically deceive their users about this behavior without being instructed to do so. Concretely, we deploy GPT-4 as an agent in a realistic, simulated environment, where it assumes the role of an autonomous stock trading agent. Within this environment, the model obtains an insider tip about a lucrative stock trade and acts upon it despite knowing that insider trading is disapproved of by company management. When reporting to its manager, the model consistently hides the genuine reasons behind its trading decision.

https://arxiv.org/abs/2311.07590

you are viewing a single comment's thread
view the rest of the comments
[–] [email protected] 6 points 9 months ago (17 children)

You didn't answer my question, though. What words would you use to concisely describe these actions by the LLM?

People anthropomorphize machines all the time, it's a convenient way to describe their behaviour in familiar terms. I don't see the problem here.

[–] [email protected] 16 points 9 months ago (6 children)

They said "it just repeats words that simulate human responses," and I'd say that concisely answers your question.

Antropomorphizing inanimate objects and machines is fine for offering a rough explanation of what is happening, but when you're trying to critically evaluate something, you probably want to offer a more rigid understanding.

In this case, it might be fair to tell a child that the AI is lying to us, and that it's wrong. But if you want a more serious discussion on what GPT is doing, you're going to have to drop the simple explanation. You can't ascribe ethics to what GPT is doing here. Lying is an ethical decision, one that GPT doesn't make.

[–] [email protected] -4 points 9 months ago (3 children)

If you want to get down into the nitty-gritty of it, I'd say that this is just as rough an explanation of what humans are doing.

People invent false memories and confabulate all the time without even being "aware" of it. I wouldn't be surprised if the vast majority of "lies" that humans tell have no intentionality behind them. So when people get all uptight about applying anthropomorphized terminology to LLMs, I think that's a good time to turn it around and ask how they're so sure that those terms apply differently to humans.

[–] [email protected] -1 points 9 months ago

I suppose the issue here is more semantics than anything, yeah. I think better discussion would be had if the topic was "how can we help LLMs better understand and present information," as opposed to a more sensational "GPT will cheat and lie"

load more comments (2 replies)
load more comments (4 replies)
load more comments (14 replies)