FlowCut Edge — NVIDIA On-Device AI

Run NVIDIA AI models on the ASUS Ascent GX10 (NVIDIA GB10 Blackwell) for FlowCut's AI features: chat, vision understanding, and video generation — all on-device.

Models

Model	Task	Size
LLaVA-NeXT 7B	Image understanding, scene description	~14 GB
Cosmos Predict 2.5 (2B)	Text-to-video, image-to-video, morph transitions	~4 GB

Quick Start

1. Deploy base service

# On the GX10:
cd ~/flowcut-edge
bash deploy.sh

This installs dependencies and downloads the text/vision models.

2. Set up Cosmos video generation (optional)

bash setup_cosmos.sh

This clones nvidia-cosmos/cosmos-predict2.5, sets up the Python environment with CUDA 13.0, and authenticates with HuggingFace. Model checkpoints (~4 GB for 2B) auto-download on first inference.

3. Start the server

bash start.sh

Server runs at http://<device-ip>:8000.

API Endpoints

Chat (OpenAI-compatible)

POST /v1/chat/completions

Automatically routes to vision or text model based on content.

Video Generation

POST /v1/video/generate/text    — Text-to-video (Cosmos)
POST /v1/video/generate/image   — Image-to-video (Cosmos)
POST /v1/video/generate/morph   — Morph transition between two frames
GET  /v1/video/download/{file}  — Download generated video

Health & Models

GET  /health      — GPU info, model status
GET  /v1/models   — List available models

Architecture

FlowCut Desktop App
        │
        ▼
  HTTP (OpenAI API)
        │
        ▼
┌─────────────────────────────┐
│   FlowCut Edge (FastAPI)    │
│   http://<gx10-ip>:8000     │
├─────────────────────────────┤
│ LLaVA-NeXT 7B    (vision)    │
│ Runway           (video/cloud)│
└─────────────────────────────┘
      ASUS Ascent GX10
     NVIDIA GB10 Blackwell
       119 GB RAM, CUDA 13

Device Info

Hardware: ASUS Ascent GX10 (NVIDIA GB10 Blackwell)
RAM: 119 GB unified memory
Storage: 916 GB NVMe
OS: Ubuntu 24.04 (aarch64)
CUDA: 13.0 / Driver 580.123

Environment Variables

Variable	Default	Description
`VISION_MODEL`	`llava-hf/llava-v1.6-mistral-7b-hf`	HF vision model ID
`DEVICE`	auto-detect	`cuda` or `cpu`

FlowCut Integration

In FlowCut settings, set:

AI Model: llava-v1.6
NVIDIA Edge API URL: http://<device-ip>:8000/v1

The app will automatically use the edge device for all AI operations, falling back to cloud providers if the device is unreachable.

Name		Name	Last commit message	Last commit date
Latest commit History 26 Commits
api		api
scripts		scripts
.gitignore		.gitignore
README.md		README.md
deploy.sh		deploy.sh
requirements.txt		requirements.txt
setup.sh		setup.sh
start.sh		start.sh

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

FlowCut Edge — NVIDIA On-Device AI

Models

Quick Start

1. Deploy base service

2. Set up Cosmos video generation (optional)

3. Start the server

API Endpoints

Chat (OpenAI-compatible)

Video Generation

Health & Models

Architecture

Device Info

Environment Variables

FlowCut Integration

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

FlowCut Edge — NVIDIA On-Device AI

Models

Quick Start

1. Deploy base service

2. Set up Cosmos video generation (optional)

3. Start the server

API Endpoints

Chat (OpenAI-compatible)

Video Generation

Health & Models

Architecture

Device Info

Environment Variables

FlowCut Integration

About

Resources

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

Packages