Twitter/X

On 2026-08-04 @literalbanana reports that ChatGPT Sol 5.6 Extra High…

Brief

The posts juxtapose a recent jump in machine-assisted plagiarism detection with structural limits of the existing industry. As of 2026-08-04, @literalbanana says ChatGPT Sol 5.6 Extra High located specific papers missed during a January 2024 manual review of Neri Oxman material, and claims previously effective evasion tactics like "tortured phrases" and patchwriting are now less reliable. Max Spero traces why large-scale detection hasn’t been done: Turnitin (founded 1998) secured exclusive Crossref access to power iThenticate, an N-gram similarity-scoring tool that most journals and universities use pre-submission. Because institutions can pre-check and revise manuscripts and Turnitin’s business model prioritizes protecting customer reputations, neither broad rescans nor disruptive algorithm changes are likely, and paywalled Crossref data raises the cost barrier for new entrants.

Why it matters

On 2026-08-04 @literalbanana reports that ChatGPT Sol 5.6 Extra High (state-of-the-art) recently found papers and specific sources they couldn't locate with manual methods during a January 2024 Neri Oxman review, though it required substantial compute per hit.

Key details

  • "Tortured phrases" and reasonably sneaky patchwriting/paraphrasing were viable ways to evade machine plagiarism detection until very recently, per @literalbanana.
  • Turnitin (founded 1998) and Crossref are the two primary players: Crossref is the DOI registry for scholarly publishing, and Turnitin negotiated exclusive access to Crossref data to power iThenticate, making iThenticate the dominant research-integrity plagiarism tool.
  • iThenticate is an N-gram matching system that returns similarity scores; widespread pre-submission use by universities and journals lets authors iteratively rewrite to lower scores, and Turnitin’s business incentives plus paywalled Crossref data make large-scale rescans or new competing detectors unlikely.
Source evidence

My semi-informed opinion is that machine plagiarism detection just wasn't that good until about yesterday. When I was looking at the Neri Oxman stuff in January of 2024, doing it by hand was way more effective than using machine help.

As of yesterday, state-of-the-art ChatGPT Sol 5.6 Extra High was able to find stuff I couldn't find by my usual methods, pointed at specific papers. I think it's still a lot of compute per hit but it's just recently in my opinion better than an obsessed banana.

Let's not forget that "tortured phrases" were a viable way to get around machine plagiarism detection until recently. Reasonably sneaky patchwriting and paraphrasing, much more so.

Max Spero (@maxspero)

Why doesn't someone just run large-scale plagiarism detection to find more fraudulent papers? I know a little bit about this space so I'll give it a stab.

There are two primary entities to know about: Turnitin and Crossref.

Turnitin today is the world's biggest provider of plagiarism detection software. Founded in 1998, they started in higher education and worked primarily on academic integrity.

Meanwhile, in the early internet, people wanted better way to link to references or objects without the links dying so they invented the DOI (I am glossing over a lot here). Scholarly publishers wanted a way to link references across journals, and through this, Crossref was born. Crossref quickly became the definitive DOI registry for scholarly publishing.

After Turnitin had been operating for years in academic integrity, they wanted to get into research integrity. Turnitin had the tech, Crossref had the database. Turnitin made a deal with them, winning exclusive access in exchange for discounts to publishers who share their data. Their tool was called iThenticate, and it is essentially the one and only plagiarism detection tool for research integrity, because of this exclusive deal for Crossref data. iThenticate is based on an N-gram matching algorithm that produces a similarity score, and similarity scores above a certain threshold are evidence of potential plagiarism.

Today, Turnitin sells iThenticate bidirectionally. They sell to the journals, to help them confirm no plagiarism before they accept and publish a paper. This is a revenue source, but because of the way Turnitin grants discounts to every Crossref member who shares their data, basically every journal buys iThenticate. This is because the universities are the bigger fish: they have more money and there are many more of them. They want to confirm that the articles they submit don't get desk rejected or accused of plagiarism. This would be a great reputational risk to the university, and essentially the service that Turnitin provides to both sides is a large reduction in reputational risk.

So, why is there so much plagiarism that doesn't get caught?

The important thing is that basically every major institution has iThenticate, and they can run it on manuscripts before submitting. This means, if somebody actually plagiarized, and they go to submit, they would see the similarity score is higher than it should be, and go back and rewrite parts of the text until the similarity score is acceptable. In the end, it's less about preventing plagiarism and more about flagging text that would be a smoking gun.

And what is stopping someone from running large scale plagiarism detection?

Turnitin's iThenticate algorithm is largely static. If it was good enough last year, it would simply be confusing to change it for this year. It would be embarrassing to all parties involved if an already-published article was found to have plagiarism that iThenticate didn't catch last year.

But even if iThenticate did update their algorithm, their business is to protect reputation, not cause a public reckoning. A large-scale re-scan would be reputationally destructive to many of their customers. No one would be happy.

And no one else has the data to build a similar plagiarism detector at scale – the paywalled data required would be massively cost-prohibitive for any new entrant coming into the scene.

— https://nitter.net/maxspero/status/2084467012504346811#m