I built a browser extension that does this, well for posts on twitter, hackernews, reddit etc.
If you want it for all text, it would also be feasible. I use a quantized mini-LM model that runs very fast and classifies eg your whole twitter feed in a couple of seconds.
Went to the website and inserted the blogposts I wrote in the last year. I had a pretty good understanding which of my blogposts had more reworks and which ones had entire passages being generated using AI and then left as is because I was happy with them. None of my articles were over 50% according to your model but the ones that I know took me a long time to write, even though I used lots of AI in the creation of them, hit below 10%, probably because I hand-edited them a lot. Overall a nice website, thanks for sharing :)
Not GP, but I too was interested in this slopdetect-minilm-v3 model.
The model appears to be similar to MiniLM-L6 (384 dimension, 6 transformer layers) but uses RoBERTa/GPT-2 style embedding/tokenization (50265 vocab size), used for binary classification, and quantized to INT8.
My guess is that is was distilled from roberta-base, then fine-tuned on freely available pile like artem9k/ai-text-detection-pile or similar.
It's a nice model - I've just used it to create a browser extension which highlights text based on how likely it is to be LLM genereated.
Edit - a quick google search reveals ibm-granite/granite-embedding-30m-english with the same architecture: slap a binary classification head on, fine tune, job done.
My pain points with PRs where people vibe coded something is a bit different though:
- I'd like to get an idea how they prompted and developed the PR.
- I want to see if for example they just took everything the AI gave them or if they interacted with it critically
- I want to see some convincing proof that they tested it, e.g. manually. I.e. along the lines of what Simon describes here: https://simonwillison.net/2025/Dec/18/code-proven-to-work/
- I want to see an AI doing a review as well
Yep that seems to be a common sentiment! Definitely some interesting areas around capturing agent conversations in commit history that could make this possible
what's the advantage of vibetunnel? And which central server is required? Vibe tunnel still sounds like "ssh into your machine from your phone", or is there something I'm missing?