I am trying to reproduce the reported performance, but I encountered a significant discrepancy. When I directly evaluate the official LIBERO checkpoint on the LIBERO-Plus benchmark, the success rate is only 19.4%, which is far below expectations.
Are there any known differences in evaluation protocol, environment version, or checkpoint compatibility between LIBERO and LIBERO-Plus that could explain this large performance gap?
Any guidance on the correct evaluation setup or potential pitfalls would be greatly appreciated. Thank you!
I am trying to reproduce the reported performance, but I encountered a significant discrepancy. When I directly evaluate the official LIBERO checkpoint on the LIBERO-Plus benchmark, the success rate is only 19.4%, which is far below expectations.
Are there any known differences in evaluation protocol, environment version, or checkpoint compatibility between LIBERO and LIBERO-Plus that could explain this large performance gap?
Any guidance on the correct evaluation setup or potential pitfalls would be greatly appreciated. Thank you!