1 2 159

Sudeep Pillai PRO

spillai

https://people.csail.mit.edu/spillai/

AI & ML interests

Self-supervised learning, Few-shot learning, Computer Vision, Robotics

Recent Activity

posted an update 9 minutes ago

mm-ctx – fast, multimodal context for agents. LLM-based agents handle text incredibly well, but images, videos, or PDFs with visual content are hard to interpret. mm-ctx gives your CLI agent multi-modal skills. Try it interactively in Spaces: https://huggingface.co/spaces/vlm-run/mm-ctx Readme: https://vlm-run.github.io/mm/ PyPI: https://pypi.org/project/mm-ctx SKILL.md: https://github.com/vlm-run/skills/blob/main/skills/mm-cli-skill/SKILL.md mm-ctx is meant to feel familiar: the UNIX tools we already love (find/cat/grep/wc), rebuilt for file types LLMs can't read natively and designed to work with agents via the CLI. - mm grep "invoice #1234" ~/Downloads searches across PDFs and returns line-numbered matches - mm cat <document>.pdf returns a metadata description of the file - mm cat <photo>.jpg returns a caption of the photo - mm cat <video>.mp4 returns a caption of the video A few things we obsessed over: ⚡ Speed: Rust core for the hot paths 🏠 Local-first, BYO model: Uses any OpenAI-compatible endpoint: Ollama, vLLM/SGLang, LMStudio with any multimodal LLM (Gemma4, Qwen3.5, GLM-4.6V). 🔗 Composable: stdin + structured outputs 🤖 Drops into any agent via mm-cli-skills: Claude Code, Codex, Gemini CLI, OpenClaw. We’d love to hear your feedback! Especially on the CLI and what file types and workflows you would like to see next.

liked a Space 1 day ago

vlm-run/mm-ctx

liked a model 8 days ago

nvidia/nemotron-ocr-v2

View all activity

Organizations

Posts 1

Post

mm-ctx – fast, multimodal context for agents.

LLM-based agents handle text incredibly well, but images, videos, or PDFs with visual content are hard to interpret. mm-ctx gives your CLI agent multi-modal skills.

Try it interactively in Spaces: vlm-run/mm-ctx

Readme: https://vlm-run.github.io/mm/
PyPI: https://pypi.org/project/mm-ctx
SKILL.md: https://github.com/vlm-run/skills/blob/main/skills/mm-cli-skill/SKILL.md

mm-ctx is meant to feel familiar: the UNIX tools we already love (find/cat/grep/wc), rebuilt for file types LLMs can't read natively and designed to work with agents via the CLI.
- mm grep "invoice #1234" ~/Downloads searches across PDFs and returns line-numbered matches
- mm cat <document>.pdf returns a metadata description of the file
- mm cat <photo>.jpg returns a caption of the photo
- mm cat <video>.mp4 returns a caption of the video

A few things we obsessed over:
⚡ Speed: Rust core for the hot paths
🏠 Local-first, BYO model: Uses any OpenAI-compatible endpoint: Ollama, vLLM/SGLang, LMStudio with any multimodal LLM (Gemma4, Qwen3.5, GLM-4.6V).
🔗 Composable: stdin + structured outputs
🤖 Drops into any agent via mm-cli-skills: Claude Code, Codex, Gemini CLI, OpenClaw.

We’d love to hear your feedback! Especially on the CLI and what file types and workflows you would like to see next.

Collections 2

models 3

datasets 0

None public yet

Sudeep Pillai PRO

AI & ML interests

Recent Activity

Organizations

Posts 1

Collections 2

Executable Code Actions Elicit Better LLM Agents

Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Executable Code Actions Elicit Better LLM Agents

VGR: Visual Grounded Reasoning

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images

Executable Code Actions Elicit Better LLM Agents

Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Executable Code Actions Elicit Better LLM Agents

VGR: Visual Grounded Reasoning

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images

models 3

spillai/llama-3-2-11b-instruct-overfit

spillai/llama-3-2-11b-instruct-amazon-description

spillai/qwen2-7b-instruct-amazon-description

datasets 0

Sudeep Pillai PRO

AI & ML interests

Recent Activity

Organizations

Posts 1

Collections 2

models 3 Sort: Recently updated

datasets 0

models 3