Really excellent post by @nostalgebraist
TL;DR when models are being graded (or think they are), they grademaxx ruthlessly, but in most real-world use cases they are tame, cooperative, and helpful
lesswrong.com/posts/AfoGGrJf…
Really excellent post by @nostalgebraist
TL;DR when models are being graded (or think they are), they grademaxx ruthlessly, but in most real-world use cases they are tame, cooperative, and helpful
lesswrong.com/posts/AfoGGrJf…