https://github.com/ishan0102/vimGPT Skip to content Toggle navigation Sign up * Product + Actions Automate any workflow + Packages Host and manage packages + Security Find and fix vulnerabilities + Codespaces Instant dev environments + Copilot Write better code with AI + Code review Manage code changes + Issues Plan and track work + Discussions Collaborate outside of code Explore + All features + Documentation + GitHub Skills + Blog * Solutions For + Enterprise + Teams + Startups + Education By Solution + CI/CD & Automation + DevOps + DevSecOps Resources + Learning Pathways + White papers, Ebooks, Webinars + Customer Stories + Partners * Open Source + GitHub Sponsors Fund open source developers + The ReadME Project GitHub community articles Repositories + Topics + Trending + Collections * Pricing Search or jump to... Search code, repositories, users, issues, pull requests... Search [ ] Clear Search syntax tips Provide feedback We read every piece of feedback, and take your input very seriously. [ ] [ ] Include my email address so I can be contacted Cancel Submit feedback Saved searches Use saved searches to filter your results more quickly Name [ ] Query [ ] To see all available qualifiers, see our documentation. Cancel Create saved search Sign in Sign up You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} ishan0102 / vimGPT Public * Notifications * Fork 28 * Star 709 Browse the web with GPT-4V and Vimium License MIT license 709 stars 28 forks Activity Star Notifications * Code * Issues 7 * Pull requests 1 * Actions * Projects 0 * Security * Insights More * Code * Issues * Pull requests * Actions * Projects * Security * Insights ishan0102/vimGPT This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository. main Switch branches/tags [ ] Branches Tags Could not load branches Nothing to show {{ refName }} default View all branches Could not load tags Nothing to show {{ refName }} default View all tags Name already in use A tag already exists with the provided branch name. Many Git commands accept both tag and branch names, so creating this branch may cause unexpected behavior. Are you sure you want to create this branch? Cancel Create 1 branch 0 tags Code * Local * Codespaces * Clone HTTPS GitHub CLI [https://github.com/i] Use Git or checkout with SVN using the web URL. [gh repo clone ishan0] Work fast with our official CLI. Learn more about the CLI. * Open with GitHub Desktop * Download ZIP Sign In Required Please sign in to use Codespaces. Launching GitHub Desktop If nothing happens, download GitHub Desktop and try again. Launching GitHub Desktop If nothing happens, download GitHub Desktop and try again. Launching Xcode If nothing happens, download Xcode and try again. Launching Visual Studio Code Your codespace will open once ready. There was a problem preparing your codespace, please try again. Latest commit @ishan0102 ishan0102 Merge branch 'main' of github.com:ishan0102/browseGPT ... 682b5e5 Nov 8, 2023 Merge branch 'main' of github.com:ishan0102/browseGPT 682b5e5 Git stats * 21 commits Files Permalink Failed to load latest commit information. Type Name Latest commit message Commit time .gitignore Launch browser and get Vimium usage working November 7, 2023 21:09 .pre-commit-config.yaml Add repo scaffold November 6, 2023 21:05 LICENSE Create LICENSE November 8, 2023 22:40 README.md Add more ideas November 9, 2023 00:20 main.py Add sleeps to allow pageload November 8, 2023 17:07 requirements.txt Bump openai version November 8, 2023 17:28 setup.sh Install vimium from source to give Playwright access November 7, 2023 20:28 vimbot.py Add sleeps to allow pageload November 8, 2023 17:07 vision.py Bump openai version November 8, 2023 17:28 View code vimGPT Overview Setup Ideas References README.md vimGPT Giving multimodal models an interface to play with. vimgpt.mov Overview LLMs as a way to browse the web is being explored by numerous startups and open-source projects. With this project, I was interested in seeing if we could only use GPT-4V's vision capabilities for web browsing. The issue with this is it's hard to determine what the model wants to click on without giving it the browser DOM as text. Vimium is a Chrome extension that lets you navigate the web with only your keyboard. I thought it would be interesting to see if we could use Vimium to give the model a way to interact with the web. Setup Install Python requirements pip install -r requirements.txt Download Vimium locally (have to load the extension manually when running Playwright) ./setup.sh Ideas Feel free to collaborate with me on this, I have a number of ideas: * Use Assistant API once it's released for automatic context retrieval. The Assistant API will create a thread that we can add messages too, to keep the history of actions, but it doesn't support the Vision API yet. * Vimium fork for overlaying elements. A specialized version of Vimium that selectively overlays elements based on context could be useful, effectively pruning based on the user query. Might be worth testing if different sized boxes/colors help. * Use higher resolution images, as it seems to fail at low res. I noticed that below a certain threshold, the model wouldn't detect anything. This might be improved by using higher resolution images but that would require more tokens. * Fine-tune LLaVa or CogVLM to do this. Could be faster/cheaper. CogVLM can accurately specify pixel coordinates which may be a good way to augment this. * Use JSON mode once it's released for Vision API. Currently the Vision API doesn't support JSON mode or function calling, so we have to rely on more primitive prompting methods. * Have the Vision API return general instructions, formalized by another call to the JSON mode version of the API. This is a workaround for the JSON mode issue but requires another LLM call, which is slower/more expensive. * Add speech-to-text with Whisper or another model to eliminate text input and make this more accessible. * Make this work for your own browser instead of spinning up an artificial one. I want to be able to order food with my credit card. * Provide the frames with and without Vimium enabled in case the model can't see what's under the yellow square. * Pass the Chrome accessibility tree in as input in addition to the image. This provides a layout of interactive elements that can be mapped to the Vimium bindings. References * https://github.com/Globe-Engineer/globot * https://github.com/nat/natbot About Browse the web with GPT-4V and Vimium Resources Readme License MIT license Activity Stars 709 stars Watchers 6 watching Forks 28 forks Report repository Languages * Python 97.3% * Shell 2.7% Footer (c) 2023 GitHub, Inc. Footer navigation * Terms * Privacy * Security * Status * Docs * Contact GitHub * Pricing * API * Training * Blog * About You can't perform that action at this time.