Model Accent has launched — a free browser-based quiz "Which AI wrote this?": visitors are shown eight short technical answers and asked to guess which assistant wrote each — ChatGPT, Claude, Gemini, or Grok. The project is anonymous and does not publish an author or text selection methodology, and all discussion so far boils down to a single Hacker News thread. Formally, this is a couple of minutes of fun, but in essence — a popular blind test of how similarly flagship models now write.

What happened
A free quiz "Which AI wrote this?" opened on the site https://modelaccent.com. The reader is sequentially shown eight short technical answers — for example, about why fast token scoring does not speed up agentic pipelines — and asked to determine which of four AI assistants wrote each fragment. The options are ChatGPT, Claude, Gemini, and Grok, and the answer is selected with keys 1 through 4. The interface is built on Next.js: there are no pages with author information, a landing page with a project description, FAQ, or description of the text selection methodology on the site. The only public discussion is a thread on Hacker News, where at the time of the review the quiz had four points and one comment, and users share their own guessing metrics.
Context
The quiz stands apart from detectors like GPTZero: they try to programmatically guess whether a text was written by a model, while Model Accent calculates nothing and delegates the assessment to a human. In essence, this is the classic blind test mechanic: the participant does not know the layout of answers by model and relies only on the feeling of the style. Behind this is the controversial thesis of stylistic convergence: common alignment techniques and typical prompt templates push flagship models toward a similar writing style, although this is an interpretation, not a measured fact. The basic threshold for such a game is also important: with four options, a quarter of the answers are guessed randomly, so any talk of "distinguishable handwriting" only makes sense with large samples of attempts.
Why this matters for the industry
For the industry, guessing quizzes like Model Accent work as spontaneous benchmarks of stylistic homogeneity of flagship assistants: if visitors consistently cannot determine the model, it means the stylistic difference between provider answers has almost disappeared, and differentiation shifts to price, context size, and integrations. The site itself is useless for production: the project has no API, pricing, methodology, or accumulated statistics, and the value is not the platform, but the mechanic — the blind test "guess the model" is a cheap and reusable product pattern that can be embedded in your own assistant comparisons. If the format is picked up and quizzes with an open protocol and aggregation of accuracy across thousands of attempts appear, for the first time it will be possible to have a crowdsourced measure of stylistic distinguishability of flagships. Until then, the practical conclusion is one: do not choose a model by "feeling of style," but compare candidates by price, context, latency, and reliability, recording the choice with your own evals.
Why this matters for users
For the ordinary reader, the quiz gives a couple of minutes of self-test: eight short texts, four options, and an honest result — it will become clear whether you really recognize the "handwriting" of your favorite assistant or all four write like written clones. Seeing the stylistic similarity of answers with your own eyes is more useful than any review article: after this, a healthy skepticism appears toward confident statements about the "house style" of one model or another. If everything is guessed in a row, this is also a diagnosis — only of your memory for typical formulations, not proof of expertise.
What is still unknown / limitations
The project has no author, methodology, or dataset: it is not described how the texts were selected — which model versions, which prompts, what generation date, and how many candidates were rejected, so there is a possible selection of deliberately "averaged" answers, artificially increasing indistinguishability. There is also no ground truth verification: the reader must take it on faith that each answer was really generated by the stated model, which makes the experiment non-reproducible. The social signal is still microscopic — one Hacker News thread with a single comment (three points in the raw thread text), and without aggregated statistics it is impossible to understand whether real user results are below the random guessing threshold. Finally, the thesis of stylistic convergence of flagships remains an interpretation, not a measured fact.
Sources
Author
Look at AI, editorial
