OneJev
View on GitHub🚀🚀 A multimodal System One decision model that gives calibrated answers to typed questions about screens, photos, video and text in one forward pass.
OneJev is a multimodal decision model and Python serving toolkit that returns calibrated probabilities for typed questions about screenshots, photos, video, and text. It includes PyTorch and llama.cpp backends, an API, and training code for GUI-agent decision tasks.
Use Cases
Choose the next action for a GUI agent from a screenshotEstimate whether a computer-use task is completeScore task progress from screen imagesAnswer typed questions about images, video, and textServe calibrated multimodal decisions through an API
Built With
- Language
- Python
- Frameworks
- PyTorch · Hugging Face Transformers · llama.cpp · FastAPI
Tags
multimodal · vision-language · computer-use · GUI agents · decision model · calibrated probabilities · video understanding · model serving · fine-tuning