Skip to content
pando/ docs

Embedding Models for Code

Which embedding models work for the code index, what each was trained on and how they compare. For the idea, read Code search by meaning; for the setup in the Web UI, the guide Choose a model that reads code.

Keys

[Remembrances]
UseSameModel          = false          # true = the code index uses the document model
CodeEmbeddingProvider = 'ollama'       # 'ollama', 'openai', 'anthropic' or 'openai-compatible'
CodeEmbeddingModel    = 'hf.co/ggml-org/jina-embeddings-v2-base-code-Q8_0-GGUF:Q8_0'
CodeEmbeddingBaseURL  = ''             # only for 'openai-compatible' or a remote Ollama
CodeEmbeddingAPIKey   = ''             # stored encrypted

Web UI: Settings > Remembrances > Code embeddings. With UseSameModel = true, or with the code fields empty, the code index uses DocumentEmbeddingProvider and DocumentEmbeddingModel.

After changing the model, run Re-index all (Settings > Remembrances > Code indexing). Vectors made by another model are not comparable, and vectors of a different size are skipped by the search.

Models that work in Ollama

All were pulled and tested with Ollama 0.35 in October 2026. GGUF is the file format Ollama loads; hf.co/<user>/<repo>:<tag> pulls a GGUF file directly from Hugging Face.

ModelSizeVectorContextTrained onLicense
CodeRankEmbed (137M)90 MB7688192Code in Python, Java, JavaScript, PHP, Go, RubyMIT
CodeRankEmbed, less compressed146 MB7688192SameMIT
Jina Embeddings v2 Base Code (161M)172 MB7688192English and 30 programming languagesApache 2.0
Qwen3 Embedding 0.6B639 MB102432KText in 100+ human languages, and codeApache 2.0
Nomic Embed Text v1.5 (general text)274 MB7688192Ordinary textApache 2.0

The name to give Ollama and to type in Code embedding model:

ollama pull hf.co/brandtcormorant/CodeRankEmbed-Q4_K_M-GGUF:Q4_K_M         # CodeRankEmbed
ollama pull hf.co/brandtcormorant/CodeRankEmbed-Q8_0-GGUF:Q8_0             # CodeRankEmbed, less compressed
ollama pull hf.co/ggml-org/jina-embeddings-v2-base-code-Q8_0-GGUF:Q8_0     # Jina Embeddings v2 Base Code
ollama pull qwen3-embedding:0.6b                                           # Qwen3 Embedding 0.6B
ollama pull nomic-embed-text                                               # Nomic Embed Text v1.5

The 30 languages of Jina v2 Base Code: Assembly, Batchfile, C, C#, C++, CMake, CSS, Dockerfile, FORTRAN, Go, Haskell, HTML, Java, JavaScript, Julia, Lua, Makefile, Markdown, PHP, Perl, PowerShell, Python, Ruby, Rust, SQL, Scala, Shell, TypeScript, TeX and Visual Basic.

Smaller files of Jina v2 Base Code (from 83 MB) are in hf.co/second-state/jina-embeddings-v2-base-code-GGUF, with tags such as Q4_K_M.

A small comparison

This is a home-made check, not a benchmark. 48 snippets (six tasks, each written in Python, Go, TypeScript, Java, Rust, PHP, C# and shell) with meaningless function names, so the model has to read the code. For each task a plain-English question is asked and we count how many of the eight right snippets are among the first eight results. Speed is 1,500-character pieces of real source files embedded per second on one desktop graphics card; read it as a ratio between models, not as what you will get.

ModelRight answers foundWeak spotsPieces per second
Nomic Embed Text (general text)67 %Go 0 of 6, shell 1 of 6about 95
CodeRankEmbed Q4_K_M88 %shell 1 of 6about 100
CodeRankEmbed Q8_088 %shell 1 of 6about 100
Jina v2 Base Code Q8_0100 %noneabout 85
Qwen3 Embedding 0.6B100 %noneabout 25

Reading it:

  • A general text model is the weak choice for code, and the gap is largest in languages it saw little of.
  • CodeRankEmbed is the lightest. It is strong in the six languages it was trained on and weaker outside them.
  • Jina v2 Base Code covers the most languages for its size and speed.
  • Qwen3 Embedding matches it and also reads many human languages, at four times the cost in time and space.
  • The more compressed CodeRankEmbed file gave the same results as the larger one.

CodeRankEmbed is designed to receive questions with the prefix Represent this query for searching relevant code: . Pando does not add it. In the check above the prefix raised its result from 88 % to 92 %.

Models that did not work or were not tried

ModelStatus
hf.co/limcheekin/CodeRankEmbed-GGUF:Q4_K_MThe repository is no longer public on Hugging Face; a new pull fails. It was the default of the setup assistant. A copy already on your machine keeps working
hf.co/jinaai/jina-code-embeddings-0.5b-GGUFDownloads, but Ollama answers 501 Not Implemented when asked for embeddings. It needs a setting (last-token pooling) that Ollama does not expose. Non-commercial license
nomic-embed-code (7B)Not tried. From 2.8 GB to 7.5 GB: a different weight class, slow without a large graphics card
Qodo-Embed-1-1.5BNot tried. Only community GGUF files (about 1.6 GB); own license (QodoAI-Open-RAIL-M)

Related