Back to Blog
    benchmarkvector-searchvertex-aiendee

    Endee Vs Vertex AI

    Cover image for Endee Vs Vertex AI

    Endee vs Vertex AI: A Comprehensive Vector Database Benchmark

    In the rapidly evolving landscape of AI and semantic search, choosing the right vector database can significantly impact the performance and cost-efficiency of your applications. In this post, we present a detailed benchmarking comparison between Endee and Google Vertex AI's Vector Search.

    We used the open-source benchmarking tool VectorDBBench and Cohere 1M vectors (768 dimensions) dataset to evaluate both systems across key metrics: Recall, Queries Per Second (QPS), and p99 Latency.

    Notably, Endee achieves these results on a significantly smaller hardware footprint Endee (4 vCPU / 16 GB) compared to Vertex AI (16 vCPU / 60 GB), highlighting its remarkable efficiency.


    1. Recall vs Top-K

    The first test evaluates the retrieval quality (Recall) as the number of retrieved items (Top-K) increases, holding the throughput relatively constant at roughly 800 QPS.

    Hardware & Configuration:

    • Vertex AI: n1-standard-16 (16 vCPU / 60 GB) | approx_neighbors=128
    • Endee: 4 vCPU / 16 GB | m=32, ef_con=256, Precision=int16

    Recall vs Top-K

    Data Table

    Top-KVertex AI RecallEndee Recall
    30.89970.9923
    50.89320.9934
    100.88930.9918
    150.85800.9911
    300.77760.9867

    Takeaway: Even with a quarter of the vCPUs, Endee maintains an exceptional recall rate of over 98.6% across all Top-K ranges, whereas Vertex AI experiences a steady drop-off, falling to 77.7% at Top-K=30.


    2. Queries Per Second (QPS) vs Concurrency

    Next, we look at throughput. For this benchmark, we tuned the parameters to hold the recall constant at approximately 97.3% for both systems to ensure an apples-to-apples comparison of raw speed.

    Hardware & Configuration:

    • Vertex AI: n1-standard-16 | leaf_nodes_to_search=0.195
    • Endee: 4 vCPU / 16 GB | m=16, ef_con=128, ef_search=128, int16

    QPS vs Concurrency

    Data Table

    ConcurrencyVertex AI QPSEndee QPS
    2140.81661.13
    4279.661295.04
    8544.991881.23
    161079.522091.50

    Takeaway: Endee significantly outperforms Vertex AI in throughput. At a concurrency level of 16, Endee handles over 2,000 QPS on a 4-core machine, compared to Vertex AI's ~1,080 QPS on a 16-core machine.


    3. P99 Latency vs Concurrency

    Finally, we measure responsiveness. P99 latency indicates the maximum time taken for the fastest 99% of queries, making it a critical metric for real-time applications. Again, recall is held constant at ~97.3%.

    P99 Latency vs Concurrency

    Data Table

    ConcurrencyVertex AI P99 Latency (ms)Endee P99 Latency (ms)
    259.23.7
    468.73.7
    862.53.8
    1625.33.7

    Takeaway: Endee provides ultra-low and incredibly stable p99 latencies, remaining under 4 milliseconds regardless of the concurrency level. Vertex AI shows much higher and more variable latencies, peaking near 69ms.


    Conclusion

    The benchmarking data speaks for itself. Endee delivers superior recall, nearly double the maximum throughput, and dramatically lower latency, all while utilizing a fraction of the compute resources required by Vertex AI.

    Whether you're building high-speed semantic search, real-time recommendation engines, or scalable RAG pipelines, Endee offers a clear performance and cost advantage.

    Interested in testing Endee for your own workloads? Reach out to our team to get started.