Objective
Develop a fair comparison for a bounded consumer task, accounting for setup costs and different model capabilities.
What would be tested
- Use the same synthetic inputs and scoring rubric across both execution paths.
- Record cold and warm latency separately and include network time.
- Document model version, hardware, memory, token limits, and network conditions.
- Compare data transfer and failure behavior without treating speed as the only outcome.
Current observations
No benchmark results are published. A local model and a cloud model may differ in capability, so the planned comparison reports task quality alongside resource and latency measures.
Next step
Select one task and define the minimum acceptable output quality before timing either path.
How the sandbox works