Rethinking Length-Based Training: Batch Composition and Loss Normalization in Speech Token Language Models · Bharat Hunt