@simonko912 Honestly, Chinchilla-optimal or undertrained beats over-baked weights any day. At least your parameters can still learn new things. What models or architectures are you working on lately? Always cool connecting with others experimenting in this space.
Just recently released a 0.41b model with the llama arch and llama tokenizer recently, trained it a bit on fineweb and then instruct tuned on a 3gb dataset for 2 epoches. Not the best but knows basic coding, math sometimes, some instructions and etc