Qwen3.5-9B-"base" finetuned for my personal RL escapades on:

  • 40% from my pretraining set: the pile, textfiles.com, stackexchange, bluesky user response modelling data, ao3, the stack, random cybernetic control loops with attractor "goals" identified and stated before the controller's actions start
  • 40% from FineWeb
  • 20% warmup data for CLM_R, a reasoning generator and reasoning discriminator for general text completion
optimizer training steps batch size schedule lr wd max_norm peft? did i sweep for these?
bnb.optim.adamw 32-bit 297 ~80.5k tokens (T^(2/3)) 10% step warmup, linear decay to 0 2e-5 0.1 1.0 full training except rank 64 rslora a=r^0.5 on in_proj_qkv, in_proj_b, in_proj_a, q_proj, k_proj, v_proj, o_proj, out_proj, up_proj, down_proj, gate_proj - merged b4 upload no
Downloads last month
6
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for crumbs-playground/qwen3.5-9b-base-me

Finetuned
(144)
this model
Free AI Image Generator No sign-up. Instant results. Open Now