I'm conducting this "benchmark" since three years. I'm still undefeated. Pleasing models are a problem.
Link: https://i.redd.it/kjhkgod2nehh1.png
Subreddit: r/OpenAI
I'm conducting this "benchmark" since three years. I'm still undefeated. Pleasing models are a problem.
Link: https://i.redd.it/kjhkgod2nehh1.png
Subreddit: r/OpenAI