Orchestrated Multi-Model AI System

About OMMAIS

OMMAIS — the Orchestrated Multi-Model AI System — answers a question with a panel of AI models instead of one, and runs on models you control.

What OMMAIS is

Most AI chat sends your question to one very large model in someone else's data centre. OMMAIS does it differently. A first model — OMMAIS itself — reads the question and decides what it needs. For anything that deserves more than a quick answer, a planning seat writes a separate brief for three specialists, each running on a different model so they do not all make the same mistakes. The specialists can hand work to each other, a judging seat weighs where they disagree, and a final seat writes the one answer you read. Figures in the experts' work are checked by arithmetic, not by another model's opinion.

You can watch all of it happen: every seat's thinking streams into the page as it works, and nothing about how the answer was reached is hidden. The models are yours — running on your own computer through Ollama, or, if you prefer, through a cloud provider with your own key.

Open OMMAIS Chat

Getting started

There are three ways in. Pick the one that fits the device you are on.

On a computer, with your own models (recommended)

  1. Install Ollama from ollama.com — it is free, and runs on Windows, macOS and Linux.
  2. Get a model. In a terminal, ollama pull llama3.2 is a small one that runs almost anywhere; qwen3:8b is a good step up if you have 8 GB of graphics memory or more. The chat's setup also has a guide to which models fit your machine, and can pull one for you.
  3. Open OMMAIS Chat in Chrome, Edge or Firefox. It checks for Ollama straight away. The first time, Ollama has to be told to accept this site — the page shows the one line to paste for your system, with a Copy button.
  4. Enter your email address, and a passphrase if you want your conversations and files kept on this device (encrypted with it). Leave the passphrase blank and nothing is stored.
  5. Ask something. Start on Balanced. The models are chosen for you; you can change any seat under Settings.

On a phone

Android: Ollama runs on the phone itself under Termux. The chat spells out every step, each with its own Copy button — including how to keep Termux from being stopped in the background, which is the part that usually goes wrong. Small models (1B–3B) run comfortably.

iPhone and iPad: iOS cannot run Ollama, so use a cloud provider instead (below). OpenRouter has free models.

With no install at all: a cloud provider

Under Cloud providers in the chat's setup, paste an API key from OpenRouter, Groq, Google Gemini, Mistral, Hugging Face, Qwen, DeepSeek, Together, Fireworks, Cerebras, OpenAI, Anthropic or xAI. Each provider's models are listed and marked free, free tier or paid, with prices where the provider publishes them. Add the ones you want and pick them for a seat. Your prompts then go to that provider, straight from your browser — never through this site.

Fast, Balanced, Deep

Three buttons beside the message box set how much work goes into an answer.

Fast
OMMAIS answers on its own. Quickest, and plenty for a simple question.
Balanced
OMMAIS answers, and if it judges that the question needs more than it can give alone, it offers the full panel — one click to run it.
Deep
OMMAIS works out what the question needs, introduces the team it has chosen and says what happens next; then the whole panel works it and writes one answer.

What it can do

Write things you keep

Articles, stories, reports and papers open in the canvas as a formatted page and save as a Word document. Ask for Markdown or another format and that is what you get.

Build working pages

Web pages, small apps and games run live in the canvas beside the chat, and Python runs in a sealed sandbox.

Draw what it explains

Charts, tables, diagrams and interactive visuals appear right in the answer when they help.

Calculate, not guess

A built-in maths, physics and chemistry engine computes figures on your machine, so numbers are worked out rather than recalled.

Read your files

Add files or a whole folder and OMMAIS answers from them, citing what it used. The blog's archive is already indexed, so you can ask about any post.

Look things up — if you let it

With lookups switched on, experts can check Wikipedia, the week's news and other public sources. Off unless you turn it on, and each lookup is shown.

Make images

If you run a Stable Diffusion WebUI (Automatic1111, Forge, SD.Next), OMMAIS can generate images through it, on your own machine.

Talk and listen

Dictate in browsers that support it, and have answers read aloud. (In Chrome, dictation goes through Google's speech service.)

Keep your work

With a passphrase, conversations, their thinking, their files and your provider keys are saved encrypted in your own browser. Export and import them as a file.

Paste anything

A long paste becomes an attached document you can see and remove, so a pasted report is read as a whole rather than as a message.

Privacy, in brief

The only thing this site collects is the email address you enter to open the chat. Your conversations, files and keys stay in your own browser. With your own models, nothing you type leaves your computer; with a cloud provider, your prompts go to that provider; with lookups on, the short search phrases go to the services named. The site uses Google Analytics. The whole account is on the privacy page.

OMMAIS Desktop

The chat on this site is a cut-down version, running in your browser. OMMAIS Desktop is the full application for Linux: a local AI workstation that coordinates a panel of models over your own files, proposes file changes for your approval rather than making them, and ships with its own games. A rebuilt version is in development.

The blog

The OMMAIS Blog is published most days — astronomy, physics, AI, and the ideas that connect them. Browse by topic, or search the whole archive from the search page. Every post ends with a button that takes its question straight into OMMAIS Chat, and the chat can answer questions about anything the blog has covered.

The research

The Discovery Plateau Hypothesis asks whether there is a ceiling on what can be discovered, and follows the question into information theory, thermodynamics and quantum interpretation. The papers are on SSRN and indexed on Google Scholar. One of them, Transistors and Symphonies, is the argument OMMAIS is built on: that a coordinated panel of small, specialised models is a scalable alternative to one very large one.

The books

The books include The Discovery Plateau Hypothesis, the trade edition of the research; There's Nothing to Wake Up From, which follows that argument into anomaly and consciousness; and Liminal Spaces, a series of science-horror novels and a short story set in the same world.

Communities

Eric is one of the admins who help run Everything Astronomy and the Universe Main, Everything Astronomy and the Universe Satellite, Everything Science and the Universe, Worldwide Amazing Space Events / Natural Phenomena / Weather Events and Quantum Mechanics, Mathematics, Astrophysics, Group of Physicist on Facebook. Green Vine Station, on herbalism and herbal medicine, is his own page.

About Eric

I write fiction, publish research, and build software. The three look unrelated and are not: all of them are about what finite systems can and cannot do.

I have worked on computers since the DOS era. My first machine was an Apple IIe and my first language was BASIC. I studied information technology in college, and my working expertise is in applied AI, operating systems, and hardware — the practical end, where a thing has to actually run on the machine in front of you.

That constraint is the throughline. OMMAIS exists because the industry's answer to every problem is a larger model in someone else's data centre, and I wanted to know how much could be done locally instead. The answer turned out to be most of it.

I describe myself as a natural philosopher, which is a polite way of saying I work across fields that do not usually talk to each other. My preprints are on SSRN and indexed on Google Scholar, and I have also been published in peer-reviewed journals, in lower-impact venues rather than the flagship ones. I mention it because summaries of my work sometimes state that none of it has been peer reviewed, which is not accurate.

I am equally happy to say what the work does not establish. The DPH papers label their own claims, and the labels differ: some of it is argued from established bounds, some is a hypothesis with quantitative models attached, and some is explicitly conjectural. A framework presented with more confidence than it has earned is worth less than one that is honest about its edges.

Otherwise: I play a decent game of chess — no FIDE rating, so take the claim for what it is worth. I am from Chicago and have spent most of my life downstate in Illinois.

Elsewhere

SSRN · Google Scholar · LinkedIn

Contact

eric.varney@yahoo.com