• BioMan
    link
    fedilink
    English
    arrow-up
    2
    ·
    edit-2
    10 hours ago

    Machine learning training has always presented as a power law, with exponential increases in processing required for linear increases in performance. The Chinchilla scaling laws paper and efficient compute frontier papers let them select the optimal tradeoff between number of parameters and how long to cook it, improving performance greatly by letting you predict how to best use X amount of computation for Y amount of time, setting off people spending hundreds of millions of dollars since they actually knew they could use it optimally, directly leading to the perceived massive increase in performance from 2022-2024ish as they had a one-time burst of knowing how to optimally partition computation and convince people to spend lots of money at once. Everything since that time has been exponentially diminishing returns as expected.