llms are sycophantic not just toward users, but toward their own previous work as well. if you ask an llm to review code in the same thread where it made the changes, it will almost always say that everything was fixed correctly. if you compact the context or run the review in a separate thread, the llm will almost always return far more honest and unbiased findings