Opinion
Care for a drop of Sahara Silk? Or maybe a splash of Oasis White? Each name is a would-be brand for camel milk, put forward by Google AI Overview. Not bad for a lactose-intolerant chatbot. The exercise arose from my talk with Paul Martin, the founder of Summer Land Camels, whoβs selling his milk to the world, presuming he can pick the right label. Though at this stage, with sales going well, Camel Milk nails it.
Nonetheless, it was fun watching AI Overviews deliver such gems as Nomad Nectar and Simoom Smooth. Even Humpilicious had a flair unthinkable by earlier programs β machines are getting craftier. I set the parameters β βgive me 20 names for camel milkβ β and awaited the results. While none astonished, each had a playful word-feel, that spooky knack we humans own to conjure crisp phrasing.
It stands to reason then that AI is up to cryptic thinking. If Overview and every other virtual brain can tackle camel milk, then surely decoding puns and anagrams is in reach. Tom Hawking explored that premise for Gizmodo, a tech-science website. Based in New York, Tom is an Australian journalist with a love of twisted clues.
Mine in particular. Explaining why he picked five of my Friday clues from βthe most perverse and human of puzzles, the crypticβ, a chance to test three LLMs (Large Language Models): ChatGPT, Claude Sonnet 5 and Oreate AI. The toughest clue read: January 15, 2000? (9) The wordplay relies on Roman numerals, and our home calendar, pointing to MIDSUMMER (where MM β or 2000 β mirrors the dateβs seasonal position.) In short, the chatbots struggled. Claude imploded; Oreate entered a fugue state, while ChatGPT claimed the answer was HEWITT WON.
βA terrible answer,β Tom wrote. βHEWITT WON is two words for a start. Yet, ChatGPT was so confident of this wrong answer!β That last remark echoed an experiment a friend did last week. A cryptic fan, Stuart McArthur challenged AI with his own clue: Bill Gates driving almost erratically (11) Claude replied with BILLIONAIRE, claiming βthe clincher is the right letter tally, and itβs an obviously witty description of Bill Gates himself.β But Claude was wrong, as Stuart explained, showing how the anagram (mixing GATES + DRIVIN/g) yields ADVERTISING, a synonym of bill.
I know, a mate helping a chatbot to improve, throwing my puzzle career under the bus. Though Claude held fast to its wrongness, conceding, βYes, your answer parses much better.β Dear Claude, it parses better because itβs right, unlike your billionaire which ignores the cryptic mechanics.
Both tests, plus the camel milk brainstorm, tell us where we stand in this new era. If the βthinking fieldβ is open, with no wrong answers, then AI can roam to fetch ideas, brand names, thought-starters. Unlike the two clue fails, where Tom and Stuart put Claude and other clankers through their paces. Both tests sought a unique answer, and every bot overlooked the mapβs X to find foolβs gold instead.
Far worse than foolβs gold, a self-convinced eureka, the sort to delude the unsuspecting user, where an imaginary Lleyton Hewitt may win every match. Where the flat-earth disciple says gravity is phony because here is my truth. Trust me.