A real-time voice AI that can hear, see, understand, and control your computer โ on any OS. Supports Windows, macOS, and Linux. Built on the Gemini Live API for native audio streaming, delivering zero subscriptions and total digital autonomy.
MARK LII is the personalization release: the assistant that becomes yours. Pick the voice it speaks with, tune the colour of the whole HUD, and watch it power on with a boot chime and a swelling animation like a machine coming to life. The interface now breathes with you too โ the waveform and the arc-reactor core pulse to your real voice while you speak and to JARVIS's own voice while it answers.
All of that sits on the Mark LI foundation: a plugin engine you extend without ever touching the core, native audio that hears the emotion in your voice, knows when you're not talking to it, and can hold one conversation for hours.
It's not just an assistant โ it's an extension of your digital life.
| Feature | Description |
|---|---|
| ๐๏ธ Voice Picker | Choose from 5 native Gemini voices and switch live from the UI โ no restart |
| ๐จ Live Theming | Recolour the entire HUD from a hue wheel or hex โ applied instantly across every panel |
| ใฐ๏ธ Reactive HUD | Waveform and reactor core pulse to real audio โ your mic while listening, JARVIS while speaking |
| ๐ง Recallable Memory | No size limit and nothing silently forgotten โ the prompt carries what fits, the rest is looked up on demand from a local search |
| ๐๏ธ Memory Panel | See every fact JARVIS has stored about you, when it learned it, and delete any of it in one click |
| โฉ๏ธ Undo | Take back what the assistant did โ files it moved, renamed, created or wrote, and settings it changed |
| Shutdown, restart and WiFi wait for a button you press โ the model cannot confirm its own irreversible actions | |
| ๐ง Audio Device Picker | Choose the microphone and speakers by name, filtered to the short list your OS shows โ and measured, so every entry actually works |
| ๐ Session Continuity | A dropped connection, a voice change or a device change no longer wipes the conversation |
| ๐งฉ Plugin System | Drop a single .py file into plugins/ โ JARVIS learns a new skill on next launch |
| ๐๏ธ Real-time Voice | Ultra-low latency conversation in any language via Gemini Live API |
| ๐ Affective Dialog | Hears the emotion in your voice and adapts its tone in response |
| ๐คซ Proactive Audio | Knows when you're not talking to it โ background chatter never triggers a reply |
| โพ๏ธ Unlimited Sessions | Sliding-window context compression โ one conversation can last for hours |
| ๐ฅ๏ธ System Control | Launch apps, adjust volume/brightness, WiFi, shortcuts, power โ all by voice |
| ๐งฉ Autonomous Tasks | High-level planning for complex multi-step goals via agent mode |
| ๐๏ธ Visual Awareness | Real-time screen capture and webcam vision piped into your main Gemini session |
| ๐ง Persistent Memory | Deeply remembers projects, preferences, and personal context across sessions |
| โจ๏ธ Hybrid Input | Seamlessly switch between keyboard typing and voice commands |
| ๐ Morning Briefing | On first boot: greets you, reads the time, recaps yesterday, and fetches live news |
| ๐ Proactive 2.0 | Time-aware, context-aware check-ins โ knows the time of day, your projects, and what you've been discussing |
| ๐๏ธ Session Memory | Summarises each conversation and mentions it naturally next morning โ consumed after use, never repeats |
| ๐๏ธโ๐จ๏ธ Background Monitoring | User-configured topic watching โ checks for new headlines once a day and alerts naturally |
| ๐ Hardware Monitoring | Continuous CPU, RAM, GPU and temperature telemetry with localized voice alerts |
| ๐ค๏ธ Weather Report | Live weather data for your city, personalized from memory |
| ๐บ๏ธ Dynamic Content Panel | Scrollable display layer beneath the HUD that renders web results, news, and search data |
| ๐ Multi-Mode Web Search | news / research / price / compare / search โ Gemini Grounded first, DDG fallback |
| โฐ Smart Reminders | OS-native scheduled notifications (Windows Task Scheduler / macOS LaunchAgent / Linux systemd) |
| Live flight price and availability lookup | |
| ๐ฎ Game Updater | Checks and triggers game updates on Steam and Epic Games on demand |
| ๐ File Processor | Read, summarize, and answer questions about local files |
| ๐ป Code Helper | Inline code review, debugging, and generation |
| ๐ Browser Control | Open URLs, navigate tabs, and interact with the browser by voice |
| ๐จ Send Message | Compose and send messages through WhatsApp, Telegram, and more |
| ๐ฌ YouTube Control | Search, play, and control YouTube playback by voice |
| ๐ฑ๏ธ Desktop Control | Taskbar, window management, and desktop-level operations |
| ๐งโ๐ป Silent Language Memory | Detects spoken language on first use โ all future sessions adapt automatically |
| ๐ฑ Remote Dashboard | Control the assistant from your phone via QR code pairing |
| โก Auto-Start on Boot | Registers with the OS startup system (registry / LaunchAgent / .desktop) |
| ๐ Clipboard Intelligence | Copy any text โ floating panel with Translate / Summarise / Explain / Fix |
| ๐ชช Assistant Customization | Change the assistant name, your name, voice, and colour from the UI โ takes effect immediately |
Mark LII is about making JARVIS feel like your own machine. Four small-but-delightful upgrades โ all universal: no hardcoded values, no bundled asset files, and no assumptions about your language or operating system.
JARVIS is no longer stuck with one voice. Open โ Customise Assistant and choose between five native Gemini voices โ Charon, Puck, Kore, Fenrir, Aoede โ each with its own character. The switch is live: the session rebuilds itself the instant you apply, so the new voice takes over without you restarting anything, and session resumption keeps your conversation going. The voice names are language-neutral, so the picker reads the same in every locale.
Drag the hue wheel (or type an exact hex code) and the whole HUD re-themes in real time โ panels, borders, the reactor core, the waveform, every button. Make it classic arc-reactor cyan, Iron Man gold, hostile red, or anything in between. Your choice is saved and restored on the next launch.
The waveform and the arc-reactor core now respond to real audio, not a random animation. While JARVIS listens, they pulse to your microphone; while JARVIS speaks, they pulse to its own voice โ louder speech, taller bars and a brighter, wider core. When the room goes quiet, everything settles back into a gentle idle ripple. It makes the assistant feel genuinely alive and connected to what's happening.
Every launch now opens with a proper boot: a ~2.4-second cinematic transform sound โ a reactor spinning up, servos locking into place, and a bright chord confirming "online" โ plays as the HUD swells up from a dim point, rings spin up, and a bright pulse sweeps outward. It's synthesized entirely in code (no sound file to ship, identical on Windows, macOS and Linux), and you can turn it on or off any time from the ๐ BOOT SOUND toggle in the โ controls. If a machine has no audio output, it simply stays silent โ never an error.
Built on Mark LI's foundation: the ๐งฉ Plugin System (extend JARVIS with a single drop-in file), ๐ Affective Dialog, ๐คซ Proactive Audio, and โพ๏ธ Unlimited Sessions are all still here and unchanged.
These four landed across Mark LII, LIII, LIV and LV at the same time, after each of those releases had already shipped. They are not what any one of those versions originally introduced; they are the floor all of them now stand on, so moving up a Mark never costs you something the one below it had.
No new dependencies. No bundled asset files. No hardcoded language, and nothing that assumes one operating system.
The store was capped at 2,200 characters โ the whole memory, not per entry โ because all of it was pasted into the system prompt on every connect, so growing the memory grew every request. When it filled, the oldest entries were deleted and one line was printed to a console nobody reads. An assistant advertised as remembering "projects, preferences and personal context" was in practice a two-page notepad that quietly forgot your sister's name after a few weeks.
Storage and prompt budget are now separate problems:
- Nothing is deleted. The cap is a runaway guard normal use never approaches, and if it is ever hit it says so in the activity log instead of on stdout.
- The prompt carries a core, not a dump. Identity in full, then the most recently updated facts, budgeted โ measured at under 1,000 characters on a memory holding 61 stored facts. That is smaller than the old whole-store cap, so sessions now connect with fewer tokens than before.
- The rest is fetched on demand. A
recall_memorytool searches the full store locally โ no network, no second model, well under a millisecond.
The part that is easy to get wrong: a model cannot look something up if it doesn't know the thing exists. So the prompt also carries an index of the keys it had no room for. Without it, "who is Ayลe?" gets "I don't know" while ayse_sister sits on disk unread. That index interleaves categories rather than sorting by recency โ sorted like the core, a memory with forty preferences pushed the one entry the index existed for off the end.
โ โ ๐ง MEMORY shows every stored fact, when it was learned, and a โ to forget it. Everything stays in memory/long_term.json on your machine.
JARVIS moves files, renames them, writes to them and changes your settings. None of that had a way back; if it misheard you, the only remedy was to fix it by hand.
Say "undo" โ in any language โ and it reverses its own last action:
| Files | move ยท rename ยท create ยท copy ยท write ยท delete ยท organize desktop |
| Settings | volume ยท brightness ยท dark mode |
Three things it deliberately does not do:
- It does not guess. Settings undo reads the current value before changing it. Where a platform won't report that value, nothing is registered โ an undo that restores a guess is worse than no undo.
- It does not hoard. Undoing a write means keeping the old contents in memory, so files over 1 MB are excluded and it says so rather than holding a 200 MB log for the session.
- It does not delete your files to undo a copy. The reverse of a copy is removing the copy; the reverse of "create a folder" is removing it only while it's still empty.
organize_desktop gets special treatment โ one command that moves dozens of files, which made it the least reversible thing the assistant could do. It journals every move and puts all of them back in one go, cleaning up the folders it created if they're still empty.
Undo costs nothing at runtime. It appends a closure to a list; nothing in it runs unless you ask.
The old gate read like this:
if action in _DANGEROUS_ACTIONS: # {"restart", "shutdown"}
confirmed = str(params.get("confirmed", "")).lower()confirmed is a tool parameter, which means the model fills it in. Nothing stopped it sending confirmed=yes on the first call and nothing checked that a human was ever involved. It was a convention, not a gate. And its coverage was two actions โ so toggle_wifi, which cuts the assistant's own connection to the Live API and therefore cannot be asked to undo itself, went through with no gate at all.
The token is now issued by the interface. Shutdown, restart and WiFi put a banner on the HUD and return immediately; the action runs only if you press CONFIRM. Nothing blocks โ JARVIS keeps talking while the banner is up โ so this is cheaper than the old gate, which burned two tool round trips on every power command.
The split between the two mechanisms is about reversibility, not about how alarming a word sounds. Anything undoable is done at once; only the genuinely irreversible asks. An assistant that checks with you before turning the volume down is one you stop talking to.
Both audio streams opened with no device argument at all, so they always took whatever the OS called "default" โ and on Windows that moves on its own the moment you plug a headset in. "JARVIS can't hear me" almost always meant "JARVIS is listening to the webcam".
โ โ ๐ง AUDIO DEVICES lets you pick the microphone and the speakers by name. Two things matter more than the dropdown:
The list is short. query_devices() returns one entry per device ร host API, not per device โ measured on an ordinary Windows machine, 41 entries for what the sound settings show as 4 microphones and 4 speakers. The same microphone appears four times, under MME, DirectSound, WASAPI and WDM-KS, with nothing to say which is which. That is not a choice, it's a quiz. The picker takes one host API per direction, drops the "Sound Mapper" and "Primary Sound Driver" pseudo-devices that just mean "default", and deduplicates. 41 โ 8.
Every entry has been measured, not assumed. The obvious approach is to pick the host API with the nicest names โ WASAPI on Windows, which in shared mode doesn't resample, so with 16 kHz in and 24 kHz out against 48 kHz hardware every open failed. Adding a rate check and moving to DirectSound passes that test on both sides, and PortAudio's DirectSound output is a silent sink: the stream opens, every write returns success in ~0 ms, and not one sample reaches the speakers.
| write(2.0 s) took | ||
|---|---|---|
| MME | 2.02 s | consumed in real time |
| DirectSound | 0.00 s | swallowed instantly |
No capability flag reports that. So the app measures it โ once per host API per direction, on a background thread at startup, using silence. Two consequences worth stating plainly:
- Each direction picks its own host API. On Windows this lands on DirectSound for the microphone and MME for the speakers โ a split no amount of reasoning would have produced.
- The probe runs in the mode the app actually ships. DirectSound input passes a callback stream and fails a blocking read; probing the wrong mode rejected a microphone that works perfectly.
Your choice is stored by name, not by index โ indices shift whenever something is plugged in. If the saved device is gone, it falls back to the system default and says so in the log rather than failing to start.
session_resumption was switched on in the config and the handle the server sent back was never read โ so every reconnect started an empty session. A dropped packet, or simply changing the voice, wiped the conversation. "Unlimited sessions" leaked through exactly this hole.
The handle is captured and replayed now. A network blip, or switching your microphone, keeps the conversation intact.
It is held in memory only, deliberately: writing it to disk would make a fresh launch continue yesterday's chat, which sounds appealing but breaks the session-summary flow โ a conversation that never ends never produces a summary, and the "yesterday we talked aboutโฆ" line in the morning briefing silently disappears. Changing the voice also starts clean on purpose, since resuming restores the server's session state and would likely bring the old voice back with it.
- The assistant could die on a log line. Status lines carry emoji and arrows (
๐ค file_controller โ Moved: a.txt โ Documents/). On a non-UTF-8 console โ cp1254 on a Turkish Windows, cp1251 on a Russian one, cp932 on a Japanese one โ printing one raisesUnicodeEncodeError, and because that print sits after the tool's owntry/except, it escaped into the receive loop and took the session down. - Every computer command paid for two model round trips.
computer_settingsmade an entire second Gemini call, inside the tool, purely to translate the request into one of its own action names โ because the declaration only said "The action to perform", so the model rarely filled it in. When that second call failed, the fallback wasdescription.lower().replace(" ", "_"), which turns the Turkish for "turn it down" intosesi_kisand straight into "Unknown action". The declaration now names all 56 actions and the rest is spelling tolerance handled locally bydifflibin microseconds. When nothing matches it suggests real action names instead of dead-ending. - An unresolvable saved audio device, or one the driver refuses to open, falls back to the system default and says so โ on both the microphone and the speakers.
- A rejected session-resumption handle is dropped after one attempt, so an expired handle can never be replayed on every retry and prevent the reconnect it exists to protect.
| Mark | Focus |
|---|---|
| XLVIII | Instant interrupt ยท parallel news ยท two-phase briefing ยท exponential backoff ยท vision cooldown |
| XLIX | Auto-start ยท clipboard intelligence ยท assistant customization |
| L | Session memory ยท background monitoring ยท proactive 2.0 ยท instant vision ยท parallel news search |
| LI | Plugin system ยท affective dialog ยท proactive audio ยท unlimited sessions |
| LII | Voice picker ยท live theming ยท reactive HUD |
| LIII+ | Plugin files: email ยท quiz mode ยท calendar ยท and more |
git clone https://github.com/FatihMakes/Mark-LII.git
cd Mark-LII
pip install -r requirements.txt
python main.py
โ ๏ธ Installation Note: Some OS-specific dependencies are not bundled inrequirements.txtto keep the repo lightweight. If you hit aModuleNotFoundError, install the missing package withpip install <module_name>.
| Requirement | Details |
|---|---|
| OS | Windows 10/11, macOS, or Linux |
| Python | 3.11 or 3.12 |
| Microphone | Required for voice interaction |
| Speakers | Required for voice replies |
| API Key | Free Gemini API key (config/api_keys.json) |
Mark LII/
โโโ main.py # Core loop โ Gemini Live session, audio I/O, live audio levels, tool dispatch
โโโ ui.py # PyQt6 HUD โ reactive waveform, boot animation, log panel, plugin manager, camera feed
โโโ setup.py # First-run configuration wizard
โโโ plugins/
โ โโโ _template.py # Copy this to write a new plugin โ one file, drop in, done
โโโ actions/
โ โโโ web_search.py # Gemini + DDG parallel search (news, research, price, compare)
โ โโโ screen_processor.py # Screen capture & webcam vision via Gemini Live
โ โโโ background_monitor.py # User-configured topic watching โ daily DDG check, no crypto
โ โโโ proactive.py # Proactive 2.0 โ time/context/rotation-aware check-ins
โ โโโ reminder.py # OS-native scheduled notifications
โ โโโ system_monitor.py # CPU / RAM / GPU / temperature telemetry
โ โโโ computer_settings.py # Volume, brightness, WiFi, power
โ โโโ computer_control.py # Keyboard shortcuts, mouse, window management
โ โโโ open_app.py # Application launcher
โ โโโ browser_control.py # Web browser control
โ โโโ file_controller.py # File system operations
โ โโโ file_processor.py # Document reading and summarization
โ โโโ send_message.py # Messaging integration
โ โโโ weather_report.py # Live weather data
โ โโโ flight_finder.py # Flight search
โ โโโ youtube_video.py # YouTube playback control
โ โโโ game_updater.py # Game update management (Steam / Epic)
โ โโโ code_helper.py # Code review and generation
โ โโโ dev_agent.py # Developer task agent
โ โโโ desktop.py # Desktop and taskbar control
โโโ memory/
โ โโโ memory_manager.py # Load/save long_term.json โ sessions, monitors, identity
โ โโโ config_manager.py # api_keys.json access โ key, OS, name, voice, colour, plugin toggles
โ โโโ long_term.json # Persistent store: identity, preferences, projects, sessions, monitors
โโโ core/
โ โโโ prompt.txt # Assistant personality and tool-routing rules
โ โโโ plugin_loader.py # Plugin engine โ discovery, validation, crash isolation
โ โโโ undo.py # One shared undo stack โ actions register how to reverse themselves
โ โโโ confirm.py # Irreversible-action gate โ the token is issued by the UI, not the model
โ โโโ audio_devices.py # Microphone / speaker list โ filtered, measured, resolved by name
โโโ config/
โโโ api_keys.json # API key, OS setting, assistant name, user name, voice, UI colour, audio devices
Personal and non-commercial use only. Licensed under Creative Commons BY-NC 4.0.
Engineered by a developer building a real-world JARVIS-style assistant. โญ Star the repository to support the journey to Mark 100.
| Platform | Link |
|---|---|
| YouTube | @FatihMakes |
| @fatihmakes |