MINER

Mining multimodal internal representation for efficient retrieval

  • Visual document retrieval
  • Dense retrieval
  • Internal representations

Weien Li* · Rui Song* · Zeyu Li · Haochen Liu · Gonghao Zhang · Difan Jiao · Zhenwei Tang · Bowei He · Haolun Wu · Xue Liu · Ye Yuan

Motivation

Late-interaction retrievers preserve fine-grained signals but require many vectors per page. Dense retrievers store one vector, but typically rely on the final layer alone. MINER studies whether internal representations can improve that single-vector embedding without changing its dimensionality.

Layer-wise analysis

Normalized CKA identifies layers with retrieval-relevant structure. The alignment ratio then distinguishes earlier, less-aligned representations from the final aligned regime.

Which layers preserve retrieval structure?

Normalized CKA measures whether a layer organizes pages and queries similarly to the final retrieval space.

Layers 15–36 22 of 36 layers selected at CKA = 0.60

Markers show the measured fraction of Jina’s 2,048 neurons retained at the default cutoff.

Method: retrieval-aligned layer probing and adaptive sparse fusion

MINER aligns layer-wise representations to cross-modal final-layer anchors, then combines the selected representations into a single vector with the same dimension as the original dense embedding.

NormProbe for less-aligned layers

A normalized projection probe reweights and maps earlier candidate layers into the retrieval space.

BaseProbe for aligned layers

Element-wise feature weights preserve the geometry of the last three selected layers, which are already directly aligned.

Adaptive sparse multi-layer fusion

Validation utility determines each layer’s feature retention. A learned weighted sum and global bias produce one D-dimensional embedding.

Experiments and results

We evaluate MINER with Jina-Embeddings-v4, Eager-Embed-v1, and MoCa-3B on the ViDoRe V1, V2, and V3 visual document retrieval suites.

8 / 9 backbone × suite averages improved
+4.5% largest relative average gain
42.4× smaller than Jina late interaction · ViDoRe V2
5.3× faster than Jina late interaction · ViDoRe V2

MINER improves all three backbone averages on ViDoRe V2 and eight of nine averages across the complete evaluation. The gains come from better use of the corresponding retriever’s internal representations, while the backbone remains frozen and the output remains a single dense vector.

View backbone averages
BenchmarkMetricBackboneDense baselineMINERRelative change
ViDoRe V1nDCG@5Jina84.384.4+0.1%
ViDoRe V1nDCG@5Eager82.584.8+2.8%
ViDoRe V1nDCG@5MoCa85.785.6−0.1%
ViDoRe V2nDCG@5Jina53.355.2+3.6%
ViDoRe V2nDCG@5Eager56.058.5+4.5%
ViDoRe V2nDCG@5MoCa58.359.1+1.4%
ViDoRe V3nDCG@10Jina49.049.7+1.4%
ViDoRe V3nDCG@10Eager47.648.8+2.5%
ViDoRe V3nDCG@10MoCa49.149.5+0.8%

Retrieval quality and serving cost

Because MINER still emits one D-dimensional vector, it uses the same dense index and dot-product search procedure as its backbone.

Retrieval quality against index size for dense Jina, MINER, and late interaction MINER improves average nDCG at the same 22.24 megabyte index size as dense Jina. Late interaction scores higher but uses a 943.38 megabyte index. Dense Jina: 53.3 at 22.24 MB. MINER: 55.2 at 22.24 MB. Late interaction: 57.6 at 943.38 MB.

Showing retrieval quality against index size.

Dense Jina MINER Late interaction

Full-corpus Qdrant measurements; higher nDCG and QPS are better, lower storage is better.

On Jina and ViDoRe V2, MINER matches the dense baseline’s 22.24 MB index and retains 93% of its search throughput while improving average nDCG@5 by 1.9 points. Relative to Jina late interaction, it uses 42.4× less index storage and searches 5.3× faster.

View efficiency measurements
MethodVectors per pageQPSIndex storageAverage nDCG@5
Dense Jina115.4122.24 MB53.3
MINER-Jina114.3622.24 MB55.2
Jina late interactionMany2.75943.38 MB57.6