Qwen3-Embedding-8B is a high-performance LLM available via the InferX API, ideal for scalable text generation and natural language processing. This model delivers available completely free of charge, open weights architecture. Access Qwen3-Embedding-8B via the InferX API with reliable low-latency inference.
Tokens
Tokens
Tokens
Qwen3-Embedding-8B by InferX costs Free per 1M input tokens and Free per 1M output tokens.