Instructions to use nvidia/Cosmos3-Edge-Policy-DROID with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Cosmos
How to use nvidia/Cosmos3-Edge-Policy-DROID with Cosmos:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Add model-tree figure to Description
Browse files
README.md
CHANGED
|
@@ -53,6 +53,8 @@ Cosmos3-Edge-Policy-DROID is a robot-manipulation policy in the compact 4B Cosmo
|
|
| 53 |
- **Example usage and output:** See the [Usage](https://huggingface.co/nvidia/Cosmos3-Edge-Policy-DROID#usage) / Quickstart section.
|
| 54 |
- **Hardware:** The compact 4B size is intended for efficient single-GPU deployment; see [Usage](https://huggingface.co/nvidia/Cosmos3-Edge-Policy-DROID#usage).
|
| 55 |
|
|
|
|
|
|
|
| 56 |
Cosmos3-Edge-Policy-DROID was developed by NVIDIA as a part of Cosmos3.
|
| 57 |
|
| 58 |
Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs. It serves as a foundational building block for a broad range of Physical AI applications and research spanning world understanding, world generation, simulation, and embodied policy learning.
|
|
|
|
| 53 |
- **Example usage and output:** See the [Usage](https://huggingface.co/nvidia/Cosmos3-Edge-Policy-DROID#usage) / Quickstart section.
|
| 54 |
- **Hardware:** The compact 4B size is intended for efficient single-GPU deployment; see [Usage](https://huggingface.co/nvidia/Cosmos3-Edge-Policy-DROID#usage).
|
| 55 |
|
| 56 |
+

|
| 57 |
+
|
| 58 |
Cosmos3-Edge-Policy-DROID was developed by NVIDIA as a part of Cosmos3.
|
| 59 |
|
| 60 |
Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs. It serves as a foundational building block for a broad range of Physical AI applications and research spanning world understanding, world generation, simulation, and embodied policy learning.
|