xAI logoxAI
Coding·45 minMembers

Mock-LLM Inference Engine — Dynamic Batching

Members only

Implement a simplified inference-engine loop on top of a `model(batch) → next_tokens` API. The bar is dynamic batching: as sequences finish (max-tokens or stop-token), free their slots and fill them w...

MLE
Infra Eng
RS
ml-infra
inference
scheduling
transformer
hard
Frequency
Single report
Last asked
2025-12-14
Stage
phone-screen · tech-screen

Log in to continue reading the full content

Comments

Sign in to join the discussion
Loading...