filipe is the only tool that reads one large, complex prompt, understands it, and fans it out across every engine you own — your local GPU, your Ollama subscription, and free API tiers — then assembles it all back into one answer.
Every AI tool you use today has the same three leaks. filipe was built to plug all three at once.
Your agent re-reads the same files, re-analyzes the same context, and re-answers the same questions — every single time. That's tokens spent on things you already did.
One model, one queue, one bottleneck. Your whole prompt sits behind a single engine, even when other engines you already pay for are sitting idle.
No memory of your files, your style, or your past answers. Every session starts from zero — and you pay the full context price every time.
filipe turns one expensive, forgetful, serial prompt into a parallel, remembering, cost-aware pipeline — without you changing how you work.
filipe reads your whole request and understands what it's really asking. Then it hands each part to the engine best suited for it — local, cloud, or free — all at once.
PARALLEL BY DEFAULTEvery file you touch, every answer you keep — filipe remembers it. Next time it already knows the context, so it never re-reads, re-analyzes, or redoes what it already did.
THE TOKEN SAVERfilipe tracks your typos and learns your vocabulary, so your prompts get cleaner and more precise over time. It adapts to you — not the other way around.
GETS SMARTER DAILYYour work runs on your own GPU. Only the initial orchestration step touches the cloud — the heavy lifting stays on your hardware, by default.
YOUR HARDWARE, YOUR DATAfilipe itself costs nothing. Bring your own keys — Ollama, OpenRouter, NVIDIA, or any paid API you already pay for — and filipe makes them all work harder.
BRING YOUR OWN KEYSAn optional enhancement will make the free tool even more capable — a shared cache of skills and MCP servers, prompt rewriting, and cached AI-content search.
COMING SOONRun filipe and it boots every engine your setup unlocks — your local GPU, your Ollama subscription, and your free tiers.
filipe reads your whole prompt and figures out what it's really asking — the intent, the files involved, the shape of the answer.
Each part goes to the engine best suited for it — local for private work, free for the rest. All in parallel.
filipe assembles every stream into one answer — so you get the best of every engine, without the mess.
filipe doesn't pick a single model. It hands your prompt to every engine you have access to at once, and lets them race.
Your subscription's hosted model orchestrates the run — one call to understand your prompt and split it across every engine. The heavy lifting stays local and free.
Your own rig joins the race. Private, offline-capable answers with zero latency — and the core of filipe's token savings.
Free endpoints contribute their share at zero cost, stretching every prompt further without touching your budget.
This is where the tokens go to die. filipe keeps a living memory of your files, your past answers, and how you type. Every session starts from what you already know — not from zero.
Every file you touch is remembered. Next time, filipe already has the context — no re-reading, no re-analyzing.
Past answers become context. filipe builds on what it already did instead of starting over.
Your typos and vocabulary are tracked, so your prompts get cleaner and more precise over time.
filipe is free — you bring your own keys and your own GPU. An optional Cloudflare enhancement makes the free tool even more powerful.
filipe is free forever. It adds value by wiring together the tools you already pay for — so they work harder together than apart.
An optional enhancement, powered by Cloudflare, will supercharge the free tool with shared intelligence — so it gets faster and smarter the more it's used.
filipe is designed for power users with a local GPU and a few API keys. That's it.
filipe runs on your own hardware. A local NVIDIA GPU is required — it's the heart of the token savings.
REQUIREDYour Ollama account drives the orchestration. Add OpenRouter and/or NVIDIA keys for even more free capacity.
REQUIREDfilipe runs on Windows now. Linux and macOS support are on the roadmap.
LINUX & MAC SOONfilipe's first-run wizard walks you through everything — GPU mode, installs, and your keys.
On first run, the setup wizard starts automatically — no config.json, no manual setup.
One GPU for inference and one for embedding, both for inference, or both for embedding.
SurrealDB, Qdrant, OmniRoute, and your embedding model — each consent-gated and best-effort.
Ollama, OpenRouter, NVIDIA, or any paid API. filipe makes them all work together.
By running every engine together and remembering your work, filipe turns a single prompt into a race you always win.
Yes. filipe itself costs nothing. You bring your own keys — Ollama, OpenRouter, NVIDIA, or any paid API you already pay for — and filipe makes them all work harder together. The optional Cloudflare boost is exactly that: optional.
Yes. A local NVIDIA GPU is required — it's the heart of filipe's token savings and privacy. filipe runs on Windows today, with Linux and macOS support on the roadmap.
Two ways. First, it remembers your files and past answers, so it never re-reads or redoes work you've already done. Second, it routes each part of your prompt to the cheapest engine suited for it — so you're not paying premium cloud rates for work your local GPU or a free tier can handle.
Local-first by default. Your work runs on your own GPU. The only cloud touch is the initial orchestration step that understands your prompt — the heavy lifting stays on your hardware.
An optional enhancement, coming soon, that will make the free tool more capable — a shared cache of skills and MCP servers, prompt rewriting, and cached AI-content search. It's powered by Cloudflare and is entirely optional; the free core works without it.
If you already use AI tools and pay for tokens, filipe is for you. You don't need to understand how it works — you just run one command, and it handles the rest. The value is simple: fewer tokens, less time, one answer.
Run one prompt across every engine you own — and let filipe remember the rest.