from Hacker News

LLM Drifts: How Is ChatGPT’s Behavior Changing over Time?

by vicnov on 7/19/23, 5:17 PM with 2 comments

by vicnov on 7/19/23, 5:17 PM
"What are the main findings? In a nutshell, there are many interesting performance shifts over time. For example, GPT-4 (March 2023) was very good at identifying prime numbers (accuracy 97.6%) but GPT-4 (June 2023) was very poor on these same questions (accuracy 2.4%). Interestingly GPT-3.5 (June 2023) was much better than GPT-3.5 (March 2023) in this task. We hope releasing the datasets and generations can help the community to understand how LLM services drift better."
by jamesmurdza on 7/19/23, 5:33 PM
How is it possible to for the success rate to go from 98% to 2%? What is the author's explanation for this?