Hello, thank you very much for open-sourcing the code for this interesting and relevant benchmark!
I was wondering if you were willing to also open-source the trajectories of the models that you benchmarked in your paper (see screenshot)?
This would be very helpful for comparing against new models, and provide useful error analysis. Thank you!
Hello, thank you very much for open-sourcing the code for this interesting and relevant benchmark!
I was wondering if you were willing to also open-source the trajectories of the models that you benchmarked in your paper (see screenshot)?
This would be very helpful for comparing against new models, and provide useful error analysis. Thank you!