filipe · the token-saving orchestration CLI

One prompt. Every engine. Zero wasted tokens.

filipe is the only tool that reads one large, complex prompt, understands it, and fans it out across every engine you own — your local GPU, your Ollama subscription, and free API tiers — then assembles it all back into one answer.

free forever bring your own keys runs on your local GPU
The problem

You're paying for tokens you don't need to.

Every AI tool you use today has the same three leaks. filipe was built to plug all three at once.

🔁

It redoes your work

Your agent re-reads the same files, re-analyzes the same context, and re-answers the same questions — every single time. That's tokens spent on things you already did.

→ tokens burned on repetition
🚦

It waits on one queue

One model, one queue, one bottleneck. Your whole prompt sits behind a single engine, even when other engines you already pay for are sitting idle.

→ time lost to serial waiting
🧠

It forgets everything

No memory of your files, your style, or your past answers. Every session starts from zero — and you pay the full context price every time.

→ context paid for again and again
What filipe does

Everything your AI should have been doing.

filipe turns one expensive, forgetful, serial prompt into a parallel, remembering, cost-aware pipeline — without you changing how you work.

🧩

One prompt, many engines

filipe reads your whole request and understands what it's really asking. Then it hands each part to the engine best suited for it — local, cloud, or free — all at once.

PARALLEL BY DEFAULT
🧠

It remembers your work

Every file you touch, every answer you keep — filipe remembers it. Next time it already knows the context, so it never re-reads, re-analyzes, or redoes what it already did.

THE TOKEN SAVER
⌨️

It learns how you type

filipe tracks your typos and learns your vocabulary, so your prompts get cleaner and more precise over time. It adapts to you — not the other way around.

GETS SMARTER DAILY
🔒

Local-first privacy

Your work runs on your own GPU. Only the initial orchestration step touches the cloud — the heavy lifting stays on your hardware, by default.

YOUR HARDWARE, YOUR DATA
💸

Free forever

filipe itself costs nothing. Bring your own keys — Ollama, OpenRouter, NVIDIA, or any paid API you already pay for — and filipe makes them all work harder.

BRING YOUR OWN KEYS
☁️

Cloudflare-boosted

An optional enhancement will make the free tool even more capable — a shared cache of skills and MCP servers, prompt rewriting, and cached AI-content search.

COMING SOON
How it works

One command. Parallel everywhere.

01

Launch

Run filipe and it boots every engine your setup unlocks — your local GPU, your Ollama subscription, and your free tiers.

02

Understand

filipe reads your whole prompt and figures out what it's really asking — the intent, the files involved, the shape of the answer.

03

Route

Each part goes to the engine best suited for it — local for private work, free for the rest. All in parallel.

04

Assemble

filipe assembles every stream into one answer — so you get the best of every engine, without the mess.

filipe — session
$ filipe --engines local,cloud,free
→ understanding your prompt…
→ routing across 3 engines in parallel
● cloud ● local ● free
→ remembered 12 files · skipped 3 re-reads
● cloud done ● local done ● free done
→ assembling 3 streams…
✓ one answer assembled · tokens saved
$
The engines

Every engine you own, working at once.

filipe doesn't pick a single model. It hands your prompt to every engine you have access to at once, and lets them race.

ENGINE 01

Ollama Cloud

Your subscription's hosted model orchestrates the run — one call to understand your prompt and split it across every engine. The heavy lifting stays local and free.

roleorchestrator
ENGINE 02

Local GPU

Your own rig joins the race. Private, offline-capable answers with zero latency — and the core of filipe's token savings.

sourceyour hardware
ENGINE 03

Free API tiers

Free endpoints contribute their share at zero cost, stretching every prompt further without touching your budget.

sourcezero cost
🧠
src/app/page.tsx
src/lib/api.ts
config.json
your style · your vocab
past answers
The memory brain

filipe remembers. So you don't pay twice.

This is where the tokens go to die. filipe keeps a living memory of your files, your past answers, and how you type. Every session starts from what you already know — not from zero.

01

It knows your files

Every file you touch is remembered. Next time, filipe already has the context — no re-reading, no re-analyzing.

02

It knows your answers

Past answers become context. filipe builds on what it already did instead of starting over.

03

It knows your typing

Your typos and vocabulary are tracked, so your prompts get cleaner and more precise over time.

Free forever

Free to use. More capable with Cloudflare.

filipe is free — you bring your own keys and your own GPU. An optional Cloudflare enhancement makes the free tool even more powerful.

✓ FREE CORE

Bring your own keys

filipe is free forever. It adds value by wiring together the tools you already pay for — so they work harder together than apart.

  • Ollama + ClaudeYour Ollama subscription drives the orchestration. Required.
  • OmniRouteWired in automatically to unlock free API tiers.
  • OpenRouter / NVIDIARecommended — add one or both for more free capacity.
  • Any paid APIPlug in the paid providers you already use.
requires a local GPU · Windows today · Linux & macOS soon
☁ OPTIONAL CLOUDFLARE BOOST · COMING SOON

Make the free tool more capable

An optional enhancement, powered by Cloudflare, will supercharge the free tool with shared intelligence — so it gets faster and smarter the more it's used.

  • Skills & MCP cacheShared, deduplicated cache of skills and MCP servers — load the right tool instantly.
  • Prompt rewritingCleaner, more precise prompts before they hit your engines.
  • Cached AI-content searchWeb search that caches AI docs, model cards, and changelogs — repeat queries hit cache instantly.
the free tool stays free · the boost is optional
What you need

Built for people who already own the hardware.

filipe is designed for power users with a local GPU and a few API keys. That's it.

🎮

A local GPU

filipe runs on your own hardware. A local NVIDIA GPU is required — it's the heart of the token savings.

REQUIRED
🔑

An Ollama subscription

Your Ollama account drives the orchestration. Add OpenRouter and/or NVIDIA keys for even more free capacity.

REQUIRED
🖥️

Windows today

filipe runs on Windows now. Linux and macOS support are on the roadmap.

LINUX & MAC SOON
Get started

Up and running in one command.

filipe's first-run wizard walks you through everything — GPU mode, installs, and your keys.

filipe setup — first run
$ filipe
→ no config found · running setup…
→ detected 2 NVIDIA GPUs
→ GPU mode: one each (GPU0 inference · GPU1 embedding)
→ install SurrealDB? y
→ install Qdrant? y
→ install OmniRoute? y
→ pull embedding model? y
→ obsidian folder: ~/.filipe/obsidian
✓ config.json written
$ filipe doctor → all systems ready
$
01

Run filipe

On first run, the setup wizard starts automatically — no config.json, no manual setup.

02

Choose your GPU mode

One GPU for inference and one for embedding, both for inference, or both for embedding.

03

Install what you need

SurrealDB, Qdrant, OmniRoute, and your embedding model — each consent-gated and best-effort.

04

Add your keys

Ollama, OpenRouter, NVIDIA, or any paid API. filipe makes them all work together.

Why filipe

Speed, privacy, and cost — at once.

By running every engine together and remembering your work, filipe turns a single prompt into a race you always win.

engines working at once
0
single-queue waiting
100%
local-first by default
$0
free-tier contribution
FAQ

Questions, answered.

Yes. filipe itself costs nothing. You bring your own keys — Ollama, OpenRouter, NVIDIA, or any paid API you already pay for — and filipe makes them all work harder together. The optional Cloudflare boost is exactly that: optional.

Yes. A local NVIDIA GPU is required — it's the heart of filipe's token savings and privacy. filipe runs on Windows today, with Linux and macOS support on the roadmap.

Two ways. First, it remembers your files and past answers, so it never re-reads or redoes work you've already done. Second, it routes each part of your prompt to the cheapest engine suited for it — so you're not paying premium cloud rates for work your local GPU or a free tier can handle.

Local-first by default. Your work runs on your own GPU. The only cloud touch is the initial orchestration step that understands your prompt — the heavy lifting stays on your hardware.

An optional enhancement, coming soon, that will make the free tool more capable — a shared cache of skills and MCP servers, prompt rewriting, and cached AI-content search. It's powered by Cloudflare and is entirely optional; the free core works without it.

If you already use AI tools and pay for tokens, filipe is for you. You don't need to understand how it works — you just run one command, and it handles the rest. The value is simple: fewer tokens, less time, one answer.

Stop paying for tokens you don't need.

Run one prompt across every engine you own — and let filipe remember the rest.