OpenGVLab
/

InternVL2-2B

@@ -296,22 +296,99 @@ LMDeploy is a toolkit for compressing, deploying, and serving LLM, developed by
 pip install lmdeploy
 ```
-You can run batch inference locally with the following python code:
 ```python
 from lmdeploy.vl import load_image
-from lmdeploy import ChatTemplateConfig, pipeline
 model = 'OpenGVLab/InternVL2-2B'
 system_prompt = '我是书生·万象，英文名是InternVL，是由上海人工智能实验室及多家合作单位联合开发的多模态基础模型。人工智能实验室致力于原始技术创新，开源开放，共享共创，推动科技进步和产业发展。'
 image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/tests/data/tiger.jpeg')
 chat_template_config = ChatTemplateConfig('internlm2-chat')
 chat_template_config.meta_instruction = system_prompt
-pipe = pipeline(model, chat_template_config=chat_template_config)
 response = pipe(('describe this image', image))
 print(response)
 ```
 ## License
 This project is released under the MIT license, while InternLM is licensed under the Apache-2.0 license.

 pip install lmdeploy
 ```
+LMDeploy abstracts the complex inference process of multi-modal Vision-Language Models (VLM) into an easy-to-use pipeline, similar to the Large Language Model (LLM) inference pipeline.
+#### A 'Hello, world' example
 ```python
+from lmdeploy import pipeline, TurbomindEngineConfig, ChatTemplateConfig
 from lmdeploy.vl import load_image
 model = 'OpenGVLab/InternVL2-2B'
 system_prompt = '我是书生·万象，英文名是InternVL，是由上海人工智能实验室及多家合作单位联合开发的多模态基础模型。人工智能实验室致力于原始技术创新，开源开放，共享共创，推动科技进步和产业发展。'
 image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/tests/data/tiger.jpeg')
 chat_template_config = ChatTemplateConfig('internlm2-chat')
 chat_template_config.meta_instruction = system_prompt
+pipe = pipeline(model, chat_template_config=chat_template_config,
+                backend_config=TurbomindEngineConfig(session_len=8192))
 response = pipe(('describe this image', image))
 print(response)
 ```
+If `ImportError` occurs while executing this case, please install the required dependency packages as prompted.
+#### Multi-images inference
+When dealing with multiple images, you can put them all in one list. Keep in mind that multiple images will lead to a higher number of input tokens, and as a result, the size of the context window typically needs to be increased.
+```python
+from lmdeploy import pipeline, TurbomindEngineConfig, ChatTemplateConfig
+from lmdeploy.vl import load_image
+model = 'OpenGVLab/InternVL2-2B'
+system_prompt = '我是书生·万象，英文名是InternVL，是由上海人工智能实验室及多家合作单位联合开发的多模态基础模型。人工智能实验室致力于原始技术创新，开源开放，共享共创，推动科技进步和产业发展。'
+chat_template_config = ChatTemplateConfig('internlm2-chat')
+chat_template_config.meta_instruction = system_prompt
+pipe = pipeline(model, chat_template_config=chat_template_config,
+                backend_config=TurbomindEngineConfig(session_len=8192))
+image_urls=[
+    'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg',
+    'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg'
+]
+images = [load_image(img_url) for img_url in image_urls]
+response = pipe(('describe these images', images))
+print(response)
+```
+#### Batch prompts inference
+Conducting inference with batch prompts is quite straightforward; just place them within a list structure:
+```python
+from lmdeploy import pipeline, TurbomindEngineConfig, ChatTemplateConfig
+from lmdeploy.vl import load_image
+model = 'OpenGVLab/InternVL2-2B'
+system_prompt = '我是书生·万象，英文名是InternVL，是由上海人工智能实验室及多家合作单位联合开发的多模态基础模型。人工智能实验室致力于原始技术创新，开源开放，共享共创，推动科技进步和产业发展。'
+chat_template_config = ChatTemplateConfig('internlm2-chat')
+chat_template_config.meta_instruction = system_prompt
+pipe = pipeline(model, chat_template_config=chat_template_config,
+                backend_config=TurbomindEngineConfig(session_len=8192))
+image_urls=[
+    "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg",
+    "https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/det.jpg"
+]
+prompts = [('describe this image', load_image(img_url)) for img_url in image_urls]
+response = pipe(prompts)
+print(response)
+```
+#### Multi-turn conversation
+There are two ways to do the multi-turn conversations with the pipeline. One is to construct messages according to the format of OpenAI and use above introduced method, the other is to use the `pipeline.chat` interface.
+```python
+from lmdeploy import pipeline, TurbomindEngineConfig, ChatTemplateConfig
+from lmdeploy.vl import load_image
+model = 'OpenGVLab/InternVL2-2B'
+system_prompt = '我是书生·万象，英文名是InternVL，是由上海人工智能实验室及多家合作单位联合开发的多模态基础模型。人工智能实验室致力于原始技术创新，开源开放，共享共创，推动科技进步和产业发展。'
+chat_template_config = ChatTemplateConfig('internlm2-chat')
+chat_template_config.meta_instruction = system_prompt
+pipe = pipeline(model, chat_template_config=chat_template_config,
+                backend_config=TurbomindEngineConfig(session_len=8192))
+image = load_image('https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/demo/resources/human-pose.jpg')
+gen_config = GenerationConfig(top_k=40, top_p=0.8, temperature=0.8)
+sess = pipe.chat(('describe this image', image), gen_config=gen_config)
+print(sess.response.text)
+sess = pipe.chat('What is the woman doing?', session=sess, gen_config=gen_config)
+print(sess.response.text)
+```
 ## License
 This project is released under the MIT license, while InternLM is licensed under the Apache-2.0 license.