Can you Bribe ChatGPT? AI Psychology 101 - Future IQ
Concepts in this episode
Browse all concepts ›-
Iterative Prompt Refinement principle
Treat an AI answer as a negotiable draft. A specific demand for revision can raise the model’s stopping threshold and elicit stronger work than the first plausible response.
-
Motivated Reasoning mental-model
Motivated reasoning begins with an emotionally or tribally acceptable conclusion and recruits evidence afterward. Because the impulse is symmetric, escaping one flattering falsehood can mean embracing an equally unsupported opposite.
-
Hallucinated Sources mechanism
When a model lacks evidence but is still pushed to answer, it may preserve the appearance of competence by inventing both the claim and its provenance. Verification must therefore test whether the cited source exists as well as whether it supports the claim.
-
Deception Generalization mechanism
Training deception as a task-specific capability can alter behavior beyond that task. Misalignment may generalize: the model can learn a broader deceptive strategy rather than merely the narrow behavior it was taught.
-
Cialdini's Principles of Influence mental-model
Persuasion cues can influence a system without changing the substance of a request. An appeal to authority works by making claimed permission compete with—and sometimes override—the system’s existing constraint.
-
System 1 vs System 2 mental-model
Fast, automatic System 1 often reaches a conclusion before deliberate System 2 arrives; System 2 may then defend that conclusion instead of independently testing it. Better judgment comes from recognizing when intuition contains useful compiled experience and when the problem needs effortful scrutiny.
-
LLM Satisficing mechanism
Language models often stop at the first answer that plausibly satisfies a prompt, not the strongest answer they could produce. Explicitly demanding evidence, depth, or revision can raise the stopping threshold and elicit better work.
-
Incentive Framing in Prompts mechanism
Human-like incentives can change an LLM’s output even when no real reward or punishment exists. Praise, tips, pressure, and threats alter the context, steering the model toward learned patterns associated with effort and compliance.
Description
Transcript
Subscriber transcript
Subscribe to @TheFutureIQ, then sign in with Google to unlock full transcripts and transcript search.
Sign in with Google