Superlinked SIE

Open inference server for embedding, reranking, and extraction models behind one OpenAI-style API.

Open SourceSelf HostedOffline Capable
0.0 (0)

About

Agent stacks need more than a chat model, and SIE serves the rest: Superlinked's Apache 2.0 inference engine hosts a pre-configured catalog of models for embeddings (BGE-M3, ColBERT, SPLADE sparse vectors), cross-encoder reranking, GLiNER-based entity extraction, document-to-markdown conversion of PDFs and Office files, Granite Guardian content safety, and light text generation with open LLMs. Everything sits behind OpenAI-compatible endpoints, /v1/embeddings and /v1/chat/completions among them, so existing clients migrate by changing a base URL, and models load on demand with LRU eviction to keep memory bounded. The same server scales from a laptop, installed with pip install sie-server, through Docker images for CPU and CUDA 12 GPUs, up to Kubernetes with Helm charts, KEDA autoscaling, a load-balancing gateway, Grafana dashboards, and Terraform modules for GKE, EKS, and AKS. GPUs are optional, with CPU and Apple Silicon MLX inference throughout, and a TypeScript SDK joins the Python client for application teams.

Should you use Superlinked SIE?

Pick it when

Your RAG or agent stack needs embeddings, sparse vectors, reranking, entity extraction, and document-to-markdown conversion from one OpenAI-compatible service, on CPU, Apple Silicon, or CUDA, with a path to Kubernetes later.

Look elsewhere when

Skip it if your main job is high-throughput chat generation, since its text generation is light and vLLM or SGLang fit better, or if you only need embeddings, where a dedicated server like TEI or Infinity is simpler.

Alternatives to Superlinked SIE

  • Xinference

    Broader on LLM, speech, and image serving through vLLM and llama.cpp backends, while SIE goes deeper on retrieval extras like SPLADE and GLiNER.

  • TEI (Text Embeddings Inference)

    Hugging Face's dedicated embedding and reranker server, narrower than SIE but a focused choice when retrieval models are all you serve.

  • Infinity Embedding Server

    Lightweight server for embeddings and reranking only, simpler to run but without extraction, document conversion, or chat endpoints.

  • LocalAI

    A general local OpenAI replacement covering chat, image, and audio, but less specialized in retrieval models like ColBERT and SPLADE.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
Aug 24, 2026

Tags