xAICoding·45 minMembers
Mock-LLM Inference Engine — Dynamic Batching
Members only
Implement a simplified inference-engine loop on top of a `model(batch) → next_tokens` API. The bar is dynamic batching: as sequences finish (max-tokens or stop-token), free their slots and fill them w...
MLE
Infra Eng
RS
ml-infra
inference
scheduling
transformer
hard
Frequency
Single report
Last asked
2025-12-14
Stage
phone-screen · tech-screen
Log in to continue reading the full content
Comments
Sign in to join the discussion
Loading...
