i noticed that there are two separate scores for Qwen3.8-Max, one with tools and one without: https://unipat.ai/benchmarks/BabyVision
I am not sure how this is controlled, how do we replicate these runs with the provided parameters?
Especially for other models as like GPT 5.5, how do we ensure models have same tools for fair comparison and reproducibility?
i noticed that there are two separate scores for Qwen3.8-Max, one with tools and one without: https://unipat.ai/benchmarks/BabyVision
I am not sure how this is controlled, how do we replicate these runs with the provided parameters?
Especially for other models as like GPT 5.5, how do we ensure models have same tools for fair comparison and reproducibility?