v0.12.0 puts a realtime voice on the line that answers as quickly as a person and hands every real question to the model you already chose, with its history and tools. The voice can wear an animated 3D avatar, and administrators get multi-factor sign-in, groups inside groups, and a knob for how hard a model thinks. Shared chats can take replies from everyone they're shared with, and every saved model, tool, and skill keeps its earlier versions.
If you've used Voice mode, you know its rhythm. You speak and the transcript goes to your model. The whole reply gets written, and only then does text-to-speech read it back. The loop works, though a conversation through it never quite feels like talking to someone.
Open WebUI v0.12.0 changes who picks up the call. A realtime voice model now holds the conversation, and every question that needs real work goes to the model you already chose, tools and all. The voice can also wear a face. By the end, you'll know which switches turn all of this on and where each one lives. If you run Milvus, or several instances that share websocket traffic through Redis, read the Before you upgrade section first. There's plenty more in the changelog, which runs to more than forty additions and over two hundred seventy fixes.
A voice that talks back #
Voice calls get a second mode. Under Admin Settings, the Audio page now opens on a Voice calls section, where Call mode switches from Standard to Realtime. In Realtime mode, Voice mode talks through an OpenAI Realtime voice model, gpt-realtime-2.1-mini with the marin voice unless you pick another.
Setting Call mode to Realtime reveals the connection, voice, and prompt fields for the voice model
Your own model stays in the loop. The voice model keeps small talk to itself and passes anything that needs real work to the model picked for the chat.
Small talk stays with the voice model, while a real question makes the orange round trip through your own model
The conversation moves at the pace of a phone call, while the knowledge base, web search, or terminal you've set up still does the work behind it. When a tool needs approval or the model asks you something, the voice points you to the chat, because a spoken yes doesn't approve a tool. The base URL is yours to set, for any endpoint that speaks the OpenAI Realtime protocol, and calls need the Allow Call permission.
A face to go with it #
A voice call doesn't have to show an orb. With Realtime mode on, the model editor's Voice section gains two rows. Realtime Voice lets a model speak in its own voice, while the voice model itself stays the administrator's choice. Voice avatar gives the model a 3D character for its calls.
The new Skills switch under Builtin Tools, and the Realtime Voice and Voice avatar rows in the Voice section
The Configure button on the Voice avatar row opens Avatar setup, where you can upload a VRM file of up to 25 MiB. VRM is the open format VTubers and social VR apps use for 3D characters, and the editor accepts both its 0.x and 1.0 versions. The preview cycles through Idle, Listening, and Speaking, the three states a call moves between. Built-in motion keeps an avatar with no animations of its own breathing, swaying, and blinking.
A VRM 1.0 avatar loads in Avatar setup and moves through its Idle, Listening, and Speaking poses on built-in motion
You can replace each state's movement with a VRMA clip of up to 10 MiB. You can also add up to 16 named gestures, each with a short clip and a description of when it fits. The voice model reads those names and descriptions and plays a wave or a clap on its own when the moment calls for one.
A wave gesture, described so the voice model knows when to use it
During a call, the mouth moves with the loudness of the spoken answer. The avatar appears only in Realtime calls, and a model without one keeps the orb.
A second step at sign-in #
An instance that's reachable from the internet can now ask for more than a password. Turn on Require an authenticator for all users under Admin Settings > Authentication, and every account sets up an authenticator app on its next sign-in. The interface calls this multi-factor authentication. Turning the switch on signs every device out, your own included.
The new Multi-factor authentication section, with the main switch on and both bypass switches off
The next sign-in starts with a QR code. Scan it with any authenticator app, enter the six-digit code it shows, and Open WebUI hands back ten single-use recovery codes to download and keep somewhere safe.
A user signs in with a password, scans the QR code, enters the code from the authenticator app, saves the recovery codes, and lands in a chat
The setup screen (left) and the recovery codes (right), each shown once
From then on, every sign-in asks for a code after the password. Use a recovery code takes one of the saved codes instead, and each code works once.
A returning sign-in asks for the authenticator code, with a recovery code as the fallback
People can replace their authenticator or create fresh recovery codes themselves, from the Multi-factor authentication section of their Account settings. Either change takes a current code and signs every device out.
Each person manages their own authenticator and recovery codes from Account settings
Two more switches can let OAuth and trusted-header sign-ins through without the extra step, for an identity provider that already enforces its own. Someone who loses both the authenticator and the codes isn't locked out for good. Whoever runs the server can run open-webui mfa reset with the locked-out person's email address, which prints a one-time recovery token that works for 30 minutes.
The user edit dialog also gains Sign out all devices, which ends every session an account has open.
The Edit User dialog in v0.11.4 (left) and v0.12.0 (right), with the new Sign out all devices button
API keys aren't sessions, so they keep working after a sign-out and don't ask for an authenticator either. If you're closing off an account you suspect has leaked, remove its API keys too.
Reply in a shared chat #
People you share a chat with can now write in it, rather than reading a frozen copy or cloning it into chats of their own. Once a chat is shared with someone, the Share Chat dialog shows a Sharing mode setting. Switching the setting from Clone only to Allow replies lets everyone the chat is shared with keep the same conversation going. Folders shared with people or groups get the same setting, which only the owner or an administrator can change.
A chat shared with the Engineering group, with Sharing mode set to Allow replies
Messages from other people show the sender's name and profile picture. Answers stream to everyone who has the chat open, and a line above the message box shows who's typing.
The owner's view, with Sam's question on the left under Sam's name and Sam typing the next one
Each message is answered under the account of the person who sent it, with the models and permissions that person already has. The owner's system prompt, tools, and chat variables stay hidden from everyone else. Only the owner's messages change the chat's title, tags, and settings. People can continue, regenerate, or stop only the answers to their own messages.
When a chat stays on Clone only, people allowed to import chats get a Fork chat button that copies the conversation into their own chats up to any finished reply.
One row per setting #
Most of the model editor's settings now sit on one row each. A row names the setting and shows a short summary of its value, such as the first few capabilities and a count of the rest. Clicking a row opens it in place, and hovering a label with a dotted underline explains what that setting does.
Aria's editor in v0.11.4 (left) and v0.12.0 (right), in the same 1000 × 900 pixel window
On our test instance, Aria's editor in v0.11.4 measured 1,271 pixels from top to bottom, in a 1440 × 900 pixel browser window that showed 828 pixels of it. Reaching Builtin Tools and the voice settings meant scrolling. In v0.12.0, the same model's editor fits in that window, new Voice section included.
Every save keeps a version #
Saving a workspace model, tool, function, or skill no longer overwrites what was there. Every save that changes something keeps a version. A menu at the top of the editor lists the versions newest first, each with its author's picture and an optional commit message. Picking an older version and choosing Set as Production makes it live again.
Aria's version menu, with the version in use marked Production above the two it replaced
The editors for tools, functions, and skills can also compare a selected version with the current one, side by side or as a single diff. The prompt editor already kept a history and now gains the same comparison. Set as Production refuses a model version whose knowledge, tools, functions, or terminal are gone or no longer yours to use, so a rollback can't switch a model on half working.
Skills can now hold more than one file. Next to SKILL.md, a skill can carry scripts, references, and templates in a file tree inside the skill editor, and whole skills import from and export to ZIP or JSON. Administrators can also import skills straight from a GitHub repository or a link to a skill file. With Community Sharing on in Admin Settings > General, people allowed to export skills can send one to openwebui.com, files and all, from Share in its menu. A skill opened from openwebui.com loads straight into the skill editor for anyone allowed to import skills.
Share opens the Create post page on openwebui.com with the skill and its files already filled in, ready to post
Models can read a skill's files and edit any skill the person chatting is allowed to change. Each of those edits saves a new version, so Set as Production can roll one back. Anyone allowed to create workspace skills can also turn a chat into a new skill with /skills:create, which no longer needs a terminal.
The Skills switch under a model's Builtin Tools is on by default. Turning it off stops that model from looking up, reading, editing, or creating skills, while skills someone picks for a chat still apply.
Control how hard a model thinks #
Reasoning models let you trade speed for depth, and until now making that trade meant a trip into advanced parameters. Anyone who can edit a model can now attach Model controls to it, from the row just below Advanced Params in the model editor. A control has a name and a set of options, and each option carries its own parameters. One option can be the default, and the control shows as a menu or as a slider.
A Thinking control with three options, each setting Ollama's think level, shown as a slider with Medium as the default
People then pick an option from a new Model controls button beside the model selector in the chat input. The choice is saved to their account, so it follows them to other devices. The menu in chat shows only the labels, while the parameter values behind them stay in the model editor.
The model selector in v0.11.4 (left) and v0.12.0 (right), with the new Model controls button beside the model name
We set up a Thinking control on gpt-oss:20b through Ollama and asked the same arithmetic puzzle at each end of the slider:
Thinking | Tokens written | Reasoning | Answer |
|---|---|---|---|
Low | 25 | 30 characters | Right |
High | 209 | 624 characters | Right |
That's one question and one run at each setting, so read it as a demonstration rather than a benchmark.
The Thinking slider set to Low gets a one-line chain of reasoning. Slid to High, the regenerated answer reasons for a full paragraph and checks itself before replying
The button appears for people with both the Allow Chat Controls and Allow Chat Params permissions. Controls work on models from a server connection, not on pipes, direct connections, or arena models. Automations run with each control's default option rather than anyone's personal pick.
Groups inside groups #
A group can now have a parent. Members of the inner group get everything shared with the groups above it, along with those groups' permissions. A company can share a model with everyone once and still give each team its own settings.
The groups list in v0.11.4 (left) and v0.12.0 (right), where the same three groups now form a tree
The group editor marks inherited permissions, and switching off a permission locally doesn't take away one that comes from a parent. Deleting a group moves its subgroups up to its parent. Open browsers pick up changed access right away, without a reload.
Groups can also set Default models for new chats, which their members get instead of the instance-wide default. When several groups apply, the most deeply nested one wins, and a model someone picked as their own default still comes first.
Platform's settings, with Engineering as its parent and the default model it inherits from Engineering
On our test instance, we shared three models with Acme only and put Sam in Platform, two levels down. Sam saw all three models, and a new chat opened on the model Engineering set as its default.
Smaller prompts with many tools #
Every tool a model can call, built-in tools included, sends its full definition with each request. Those definitions can outweigh the conversation itself. Tool Search, marked experimental in Admin Settings > Interface, keeps any definition longer than the Deferral Threshold out of the request, and the threshold starts at 400 characters. The model gets those tools as a list of names with short descriptions, plus a search_tools tool that fetches a full definition when the model needs one. The tools in the request stay the same for the whole chat, so a provider's prompt cache keeps working.
Tool Search and its three settings, below the Context Compaction switch
On our test instance, a one-line question in a new chat with gpt-oss:20b through Ollama took 4,728 prompt tokens by Ollama's count, with every setting at its default. With Tool Search on and its defaults kept, Defer Built-in Tools included, the same question took 2,048 prompt tokens. We ran one model twice at each setting, and the saving on your instance depends on how many tools a model has and how long their definitions are. Tool Search works only with native function calling, and Always Loaded Tools keeps chosen tools, such as github_*, in the request in full.
Long tool runs also get room to finish. With Context Compaction on and native function calling, a reply that makes many tool calls in a row is now compacted between rounds once its context passes the token threshold. Each tool call stays beside its result. Before, compaction ran only before the reply started, so a long run could outgrow the model's context window.
Lighter under load #
Every instance now reads and writes JSON with orjson, from request bodies to streamed provider responses and live socket updates. In the measurements behind the change, an 18 KB streamed update encoded about 17 times faster and decoded about 3 times faster.
Deployments that spread websocket traffic across several servers through Redis also use less processor time while replies stream. Each server now skips live updates for chat rooms it has nobody connected to, where it used to decode every one.
Hybrid search on Milvus 2.5 and newer now runs inside Milvus, using its built-in BM25 keyword search next to the vectors in the same collection. Each search used to load every chunk of the collection into Open WebUI and score it there. Moving to Milvus's own search takes a one-time migration, covered under Before you upgrade.
Also in this release #
Some smaller changes are worth naming. A knowledge base now opens as a file browser, with its folders and files on the left and the selected file beside them. Switching between Preview and Indexed text shows the original document or the text the model searches, and anyone with write access can edit that text in place. Expanding an MCP server in the chat's Available Tools dialog now lists its tools with their descriptions and offers Reconnect when the server needs you to sign in again. OpenAI Realtime is also a text-to-speech engine now, reading replies aloud word for word with voices such as Marin and Cedar. The browser-based Pyodide code interpreter can create Word documents and PowerPoint decks, using libraries that ship with Open WebUI, so it also works without internet access. Messages sent in a channel while an earlier one is still sending, or while a file is still uploading, now wait in a queue above the message box. Each can be sent right away, edited, or deleted. Smaller admin work fills out the rest, including an Access entry in workspace item menus, an Unavailable view for models no connection offers, a Copy button in JSON Preview, input and output tokens on the Analytics dashboard, and separate switches for Tools, Functions, and tool servers.
Expanding an MCP server in Available Tools loads its tools, with a count and each tool's description
Two messages wait in a channel's queue while the first one's attachment is still uploading
Copy in JSON Preview copies the model's JSON as shown, edits not yet saved included
Things that work again #
Six fixes stand out, because each one ends a problem you may have been working around. Browsers that report a bare or regional language code, such as German or Dutch in Firefox, get the matching translation again, where v0.11.4 fell back to English. With tool approval set to ask, every tool a model calls in one turn now gets its own approval card, where only the first did and the rest stayed on Executing... forever. Pasting a large block of text into the chat input no longer freezes the page, because the editor now converts its content only when the content changes. Reasoning models no longer lose the space right after their thinking block, which could turn "The answer is 4." into "The answeris 4.". Read Aloud and voice calls play sound on iPhone and iPad. Scrolling the sidebar loads older chats again after you start a new chat or import chats, where the list stopped at the first 60 chats until a page refresh. The full changelog has every entry with links.
Before you upgrade #
A few things deserve your attention first:
This release includes security and access-control fixes. Not every one is spelled out at release time. Update at your earliest convenience.
This release includes database migrations. Back up first, and if you run several instances behind a load balancer, update them all together, since rolling updates aren't supported.
Back up your Milvus data before upgrading if you use Milvus. On its first start, Open WebUI copies every existing Milvus collection into a new one that holds the vectors, the text, and the keyword index together. Open WebUI finishes starting only once that copy is done. Nothing is re-embedded, but the copy needs free disk space for a second set of your Milvus data. On one 16-thread test machine, 200 GB took about two and a half to five hours, depending on chunk size. The originals are dropped only after every copy succeeds. Avoid upgrading Milvus and Open WebUI on the same day. On Milvus older than 2.5, the migration is skipped and hybrid search keeps working the way it did.
Instances that share websocket traffic through Redis should update together. Live chat updates from an updated instance don't reach one still on an older version. Setting WEBSOCKET_REDIS_ROOM_CHANNELS to false keeps the old delivery if you can't update everything at once. If the Redis user Open WebUI connects as has restricted permissions, room channels also need the PUBLISH, SUBSCRIBE, and PSUBSCRIBE permissions and access to the &socketio and &socketio#* channels. Open WebUI logs which permission is missing.
ENABLE_PLUGINS is now the master switch. Setting it to false turns off every kind of plugin, including external OpenAPI, MCP, and Open Terminal servers and personal direct connections, and it overrides ENABLE_TOOLS, ENABLE_FUNCTIONS, and ENABLE_TOOL_SERVERS. All four need a restart to take effect.
Links that start a chat no longer send it. A link with ?q= now fills in the message box and waits for you to press Enter, so a browser search shortcut that points at Open WebUI takes one more key press. No setting brings back the old behavior. Links that load a web page or a YouTube video, or that start a voice call, now ask first.
Three settings behave differently. ENABLE_ORJSON is now on, and setting it to false brings back the standard JSON encoder. New installations no longer show the built-in arena model, because ENABLE_EVALUATION_ARENA_MODELS now starts off, while existing installations keep their setting. A folder's default model for new chats is now chosen in the folder's own settings, and switching models inside one of its chats no longer changes the folder's default.
The full list is in the release notes, and upgrade instructions are in the documentation. Update the usual way, then set Call mode to Realtime and say hello.
Get involved #
- Update and try a realtime call. Ask it something only your own model's tools can answer, and listen for the moment it takes over.
- Tell us where it falls short. Bugs go to the issue tracker, and ideas and questions to Discussions, the Discord server, or r/OpenWebUI.
- Read the full changelog for every change, with commits and issues linked.
The Open WebUI Team




