From 24dc81a234a6e1901f3314eeadaa2813f2b78038 Mon Sep 17 00:00:00 2001 From: Weiyun Wang Date: Thu, 29 May 2025 11:15:06 +0000 Subject: [PATCH] Update README.md --- README.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 903a74b..88b449c 100644 --- a/README.md +++ b/README.md @@ -105,7 +105,7 @@ In this work, we use the Best-of-N evaluation strategy and employ [VisualPRM-8B] ### Multimodal Reasoning and Mathematics -![image/png](https://huggingface.co/datasets/Weiyun1025/InternVL-Performance/resolve/main/internvl3/reasoning.png) +![image/png](https://huggingface.co/datasets/OpenGVLab/VisualPRM400K-v1.1/resolve/main/visualprm-performance.png) ### OCR, Chart, and Document Understanding @@ -161,7 +161,7 @@ The evaluation results in the Figure below shows that the model with native mult As shown in the table below, models fine-tuned with MPO demonstrate superior reasoning performance across seven multimodal reasoning benchmarks compared to their counterparts without MPO. Specifically, InternVL3-78B and InternVL3-38B outperform their counterparts by 4.1 and 4.5 points, respectively. Notably, the training data used for MPO is a subset of that used for SFT, indicating that the performance improvements primarily stem from the training algorithm rather than the training data. -![image/png](https://huggingface.co/datasets/Weiyun1025/InternVL-Performance/resolve/main/internvl3/ablation-mpo.png) +![image/png](https://huggingface.co/datasets/OpenGVLab/MMPR-v1.2/resolve/main/ablation-mpo.png) ### Variable Visual Position Encoding