A friend sent me a PowerPoint last week and asked if I’d make it look nicer. I glanced at one column of numbers, and without reaching for a calculator, I could tell the analysis was off by a factor of a thousand. Not a rounding error. Three orders of magnitude. He hadn’t made it up. His AI had, in seconds, cleanly and confidently, and he’d pasted it straight into the deck and sent it to me to prettify.And here’s the part that gets me. He’s sharp, and in his own field, he’d have caught that error in a second. In his own wheelhouse, that same AI would have made him better. Researchers at Harvard and Boston Consulting Group ran the experiment that shows exactly where that line sits, and it turns out there are two invisible lines you cross every day. One decides whether the machine is brilliant or confidently wrong. The other decides whether you can even tell the difference.I’m not writing this to scare you off these tools. Let me show you both lines and how to stay on the right side of them.The Study: AI made them 40% better.Harvard and BCG took 758 of BCG’s own consultants, gave some of them GPT-4, and set them loose on realistic consulting work: writing, analysis, creative problem-solving, the kind of tasks these people do for a living. On that work, the ones using AI were remarkable. They finished 12% more tasks, worked about 25% faster, and their output came back more than 40% higher in quality.The detail I love is who gained the most. Not the stars. The consultants who’d scored in the bottom half beforehand improved by about 43%. The top half, about 17%. AI didn’t widen the gap between the best and the rest; it raised the floor. If that were the whole story, I’d tell you to go all in tomorrow.It isn’t the whole story.The Turn: 19 points worse.The same researchers slipped in one more task, built to sit just outside what the AI was good at. Same consultants, same afternoon, same tool. On this one, the people using AI were 19 percentage points less likely to get the right answer than the people working with no AI at all. The AI didn’t just fail to help. It pulled sharp people below where they’d have landed on their own.One caveat you shouldn’t skip. BCG helped design this study and graded a lot of it, and that’s a company with a real stake in AI looking good. So take the exact percentages as a 2023 snapshot on one model, not gospel. But the shape of the finding, brilliant here and quietly harmful there, keeps showing up everywhere once you know to look.Boundary One: The Jagged FrontierThe researchers gave this a name that stuck: the jagged frontier. Picture everything AI is good at as a territory with a coastline. Inside the border, it’s genuinely capable. Outside, it’s lost. The catch is that the border isn’t smooth, it’s jagged. Two jobs that feel equally easy to you can land on opposite sides, one safely inside and one just outside, and you can’t see the coastline. As the paper puts it, “outside of the frontier, AI output is inaccurate, less useful, and degrades human performance.”Here’s the task that tripped them. They asked the consultants to tell a chief executive which brand had the most room to grow, and handed over a spreadsheet that looked complete and pointed clearly at one answer, plus a folder of interviews. Buried in the interviews was the detail that flipped the picture to a different answer. The AI read the spreadsheet, confidently recommended the wrong brand, and the people leaning on it followed it off the edge.Understand why it was so sure. Underneath, a large language model is a prediction engine. It predicts the most plausible next words, and plausible is what it optimizes for, not true. There’s no little voice that says, I should check the interviews. So it hands you a confident answer, whether or not it’s right, and confident and correct look identical on the page.Boundary Two: The Knowledge LineThat’s the first line, the one in the work. The second runs right through you. These tools are great at getting you about 60% of the way there. But that gift only works if you already know the subject. If you can’t look at the output and feel, in your gut, whether it makes sense, you have no way to catch the moment it goes off the rails. My friend could have caught his. He knew enough. He just didn’t look.Say you’re not a chemist, and your last chemistry class was 40 years ago. You ask an AI some chemistry questions, and because you don’t know the field, they come out vague, and a prediction engine turns a vague question into a confident, vague-shaped answer. The tech crowd calls it a PhD in your pocket. But picture a real professor deciding you don’t know enough to know better, and handing you an answer that’s basically a joke. You’d nod along, because you can’t tell.So the two numbers from the start aren’t a contradiction. They’re the two sides of that coastline: same people, same afternoon, different side of an invisible line.The Real Danger: You stop checking.It gets worse because of what trust does to us. A separate study out of Microsoft Research and Carnegie Mellon found that the more people trusted the AI, the less critical thinking they did. The better it seems, the less you check, which is exactly backwards. One top comment I saw while researching this put it perfectly: they brought AI into my job, and now I’m busier than ever fixing everything it screws up.And this isn’t just careless people. Late last year, Deloitte refunded the Australian government part of a A$440,000 report: its AI had stuffed it with references to research papers that didn’t exist, plus a made-up quote from a court ruling. It looked right, so it shipped under the Deloitte name to a government client. Or the New York lawyer who filed a brief citing six court cases, every one fabricated by ChatGPT. When he asked whether they were real, it assured him they were. A judge fined him $5,000. And it’s a pattern: a researcher tracks these in a public database, and it’s passed 1,500 court filings with fabricated AI citations, climbing by the day. (How those keep getting past licensed attorneys with bar cards on the line is a whole story of its own. Different post, maybe.)Both camps are wrong.The loud arguments about AI both miss this. One camp says it’s coming for every job, so panic. The other says it’s a useless toy, so relax. They’re both wrong. What’s really going on is narrower: AI is good enough to trust, and confidently wrong exactly when you’re least able to catch it. The danger was never that it’s dumb. The danger is that it’s convincing.Centaur, cyborg, and who’s driving.So what do the people who use this well actually do? The researchers spotted two styles. The centaur draws a clean line down the middle of the work: this piece is yours, machine; this piece is mine, based on who’s better at it. The cyborg blends with the tool move by move, handing over a sentence, checking it, taking the pen back. Different styles, but in both, a human is always deciding where the line is.That gives you a simple test you can run on any AI answer before you ship it.The three-question testDoes this pass my own gut logic? My friend’s numbers failed that one in two seconds.Do I actually know enough about this to grade it, or am I trusting it precisely because I can’t?Does the AI sound certain here while I can’t independently check it?High confidence plus no way to verify is the edge of the cliff. When you hit it, slow down and do the part yourself.Now that 43% from the start means something sharper. AI raises your floor, but only while a human is standing on it, still doing the judging. The moment you hand over the thinking itself, you don’t rise. You just ship whatever it hands you, confident and wrong, like the deck that landed in my inbox.Hand the machine the part it’s good at: the first draft, the boilerplate, the 60%. But you still have to be in the driver’s seat for the part that’s actually yours, the judgment, the knowing enough to catch the lie. You can’t abdicate your thinking to a prediction engine. None of this makes me anti-AI. I reach for these tools every day, and they’ve raised my game. The people who win with them aren’t the ones who trust them the most. They’re the ones who know where the edges are.The Short VersionGiven GPT-4, 758 Harvard/BCG consultants were about 40% better inside the AI’s wheelhouse, and 19 points worse on one task built to sit just outside it. Same people, same afternoon.Two invisible lines matter: which tasks the AI is quietly wrong on, and whether you know the subject well enough to notice.These models optimize for plausible, not true. Even Deloitte shipped the fabrications to a government client.If this was useful, grab my book, Founders Who Finish, over at davesaunders.net. While you’re there, sign up for my newsletter, The Build. It’s part of a series on where AI genuinely helps, and where it quietly costs you.
AI Can Make You Worse: How to Not Cross the Invisible Line
Full Article
Original Source
Read the full article at Hackernoon →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.