Compatibility
Two separate questions get asked about compatibility, and they have separate, unrelated answers in Locus's actual architecture: which files can get indexed (a parsing question, decided per file format), and which AI assistants can query an already-indexed file (an MCP-protocol question, decided per client — not per file format). This page keeps them separate on purpose, rather than implying one depends on the other.
1. File formats — what actually gets parsed
locus index walks a folder and, per file, checks two things: is the extension in the indexed set, and is the file at or under a size cap (default 25 MB, LOCUS_MAX_FILE_MB to change it). A file that fails either check is skipped entirely — never partially indexed. The table below comes directly from the parser's own format list, not a general claim about what "text extraction" might handle.
| Format | Parseable | Indexed by default | Note |
|---|---|---|---|
| .md | Yes | Yes | |
| .txt | Yes | Yes | |
| .rst | Yes | Yes | |
| Yes | Yes | text layer only — no OCR | |
| .docx | Yes | Yes | paragraphs + table cells |
| .py | Yes | Yes | |
| .js | Yes | Yes | |
| .ts | Yes | Yes | |
| .json | Yes | Yes | |
| .yaml / .yml | Yes | Yes | |
| .toml | Yes | Yes | |
| .csv | Yes | Yes | |
| .html | Yes | Yes | |
| .htm | Yes | No | parser supports it; not in the default set |
| .xml | Yes | No | parser supports it; not in the default set |
| .sh | Yes | No | parser supports it; not in the default set |
| .sql | Yes | No | parser supports it; not in the default set |
| .doc (legacy binary Word) | No | — | |
| .xlsx / .xls | No | — | |
| .pptx / .ppt | No | — | |
| .epub | No | — | |
| .rtf | No | — | |
| images (.png, .jpg, …) | No | — | no OCR |
| audio / video | No | — |
A handful of formats the parser supports (.htm, .xml, .sh, .sql) aren't walked by default — add them with LOCUS_EXTENSIONS (comma-separated, with or without the leading dot) if you want them included. PDFs are extracted as their text layer only: a scanned or image-only PDF with no embedded text returns empty content, since Locus does text extraction, not OCR. Symlinks are never followed, hidden files/folders are skipped, and common noise directories (.git, node_modules, .venv, dist, build, __pycache__, and similar) are always excluded — the same rules whether you're on the free local engine or the hosted trial, since it's one indexing codebase either way.
Languages
The free local mode is optimized for English; managed mode uses a multilingual model. We have not yet measured search quality in other languages. Non-English files are still parsed and indexed the same way; what changes is how well search finds them. The local default is an English embedding model, so expect non-English files to be found less reliably.
2. AI clients — what can query an indexed file
Once a file is indexed, Locus hands back the same thing whichever client asks: chunk text plus a file path, from locus_search and locus_read_file. The client never parses the original file itself — Locus already did that at index time — so a client's ability to use a result has nothing to do with whether that result came from a .md file or a .pdf. A client that can talk to Locus at all can query any indexed format equally; there is no per-format client restriction anywhere in this architecture.
What does vary per client is how far we have tested it ourselves. Each client carries one of four labels, the same ones used on the full connection guide, and each label rests on a test record we keep:
How far we have tested each client
- ✓ Tested: full lifecycle(0)
- Tested, plus edits, deletions, key revocation, reconnecting, the computer going offline, and removal.
- ✓ Tested(3)
- We asked real questions through this client and saved the answers, which cited at least two files.
- ConnectsNot yet verified(11)
- We saw this client connect to a real Locus server. We have not yet recorded it answering questions.
- DocumentedNot yet verified(15)
- Setup steps written from the vendor's own docs. We have not yet finished a connection check with this client.
We only say Locus "works with" a client once it is Tested. Documented and Connects entries are labelled "Not yet verified": we have not yet recorded them returning cited answers from Locus. They are listed so you can try them.
Tested (3)
Connects (11)
Not yet verified
Documented (15)
Not yet verified
- GitHub Copilot (VS Code, JetBrains, Visual Studio)
- Antigravity
- Kimi Code CLI
- ZCode (GLM)
- LM Studio
- 5ire
- Raycast AI
- Kilo Code
- Open WebUI
- AnythingLLM
- Goose (Block / Agentic AI Foundation)
- Claude.ai (web app)
- ChatGPT (general chat, not Codex)
- Gemini Enterprise (Business Edition)
- Microsoft Copilot (via Copilot Studio)
Reading the two axes together
In short: format decides if a file is in the index at all; client decides how you ask questions about what's already in there. A Tested client asking about a .docx file and a Documented client asking about a .py file go through the exact same retrieval path once indexing is done — neither axis constrains the other. Full per-client setup steps live on /docs/connect; what's free versus trial versus paid to run any of this lives on /free, and how Locus fits alongside the note app, cloud storage, and AI assistant you already use is on /complements.
This page is generated from the same data the connection guide uses and from the parser source directly — if either changes, this page is expected to be updated in the same change, not left to drift.