NV-DINOv2 is a visual foundation model that generates vector embeddings for the input image.
Grounding dino is an open vocabulary zero-shot object detection model.