Scale what the model readslonger context → more intelligence
Intelligence begins with capacity: a model that can hold more of the world in one context is simply smarter. So I worked where long context is actually made: schedule the context window during pretraining (SkyLadder), and repair the numerics that silently break RoPE at long range (AnchorAttention).
SkyLadder: Better & Faster Pretraining via Context Window Scheduling
Short→long context scheduling: up to +3.7% on benchmarks with up to 22% faster pretraining.