If several small questions can be answered independently, sending each one as its own request adds overhead. This little Python experiment puts multiple queries into one prompt, labels each with an ID, then parses the model's ID:answer lines back into a mapping.
It targets an OpenAI-compatible chat completions endpoint, so the endpoint can be configured instead of hardcoding a single provider. Batching can reduce request overhead for this kind of workload, though the model still has to produce a well-formed answer for every ID.
Code: LLM-Batching.