Search Cluster

Our turn-key solution for your identity records. Your ultimate power Enterprise PII Platform.

PII Infrastructure

Built for Enterprise Scale

Search Cluster is a turn-key solution designed to store, search, and manage person records at enterprise scale. Built on proven cluster technology with redundancy, encryption, and semantic matching capabilities.

99.99%
Uptime SLA
AES-256
Encryption
10B+
Records Capacity
Architecture

System Overview

Your ApplicationREST API ClientYour ApplicationREST API ClientAPI GatewayAuthentication • Rate Limiting • RoutingCluster Node 1• Encrypted Storage• Semantic Search• Record MatchingCluster Node 2• Encrypted Storage• Semantic Search• Record MatchingCluster Node N• Encrypted Storage• Semantic Search• Record MatchingContinuous BackupOff-site • Point-in-time RecoveryHTTPS/TLSHTTPS/TLSLoad BalancedLoad Balanced

Encrypted End-to-End

TLS transport encryption and AES-256 storage encryption

Distributed Cluster

Automatic failover and load balancing across nodes

Continuous Backup

Off-site backups with point-in-time recovery

Use Cases

One cluster, many workflows

A single index can simultaneously support multiple use cases — automated pipelines and human-facing search side by side.

End-to-end automated processing

No human in the loop — API calls handled by defined rules. Typical patterns:

  • Searches by unique identifier (SSN, email, phone) returning matching records.
  • Upsert requests that update existing records or insert new ones based on matches.
  • Copy search requests comparing complete records against indexed data using predefined scoring rules.

Human-assisted processing

Autocomplete and full-text search designed for human interaction — users see and assess the results, then decide what to do next. Same index, different consumption pattern.

Semantic Record Matching

Match people, not strings

A two-phase approach designed for person records, CRM systems, and customer databases — well beyond traditional database search.

1

Recall

SearchCluster establishes indices that accommodate phonetic variations and simplifications. The query is transformed internally, identifying candidate matches — records sharing enough commonalities to warrant deeper examination.

2

Match

Candidate records are retrieved, decrypted, then compared entity-by-entity and attribute-by-attribute using the Matcher component — producing detailed match results delivered back to the requesting system.

Schema

Your data, your structure

Flat, relational, or deeply nested — SearchCluster ingests data in the shape you already have.

Tabular

Flat structures with columns and rows — typically from SQL databases or CSV feeds.

Relational

Multiple sets of tabular data linked through keys and IDs.

Document

Nested objects and arrays representing naturally structured information.

Index schema

SearchCluster builds an internal schema organising entities with semantic data types:

Fields

Typed elements — GIVENNAME, SURNAME, RECORDID.

Entities

Collections of fields — PERSON, ADDRESS.

Multiplicity

Fields and structures support multiple values via CSV or JSON arrays.

Record-Level Encryption

Protect the people behind the data — as well as your company — by preventing data theft.

Every Optimaize SearchCluster stores the data for each record not in plain text, but in encrypted form.

Compare a traditional database on the left, such as a relational database table, with a SearchCluster on the right — drag the slider:

Employee IDFull NameJob TitleDepartmentBusiness UnitGenderEthnicity
E02002Kai LeControls EngineerEngineeringManufacturingMaleAsian
E02003Robert PatelAnalystSalesCorporateMaleAsian
E02004Cameron LoNetwork AdministratorITResearch & DevelopmentMaleAsian
E02005Harper CastilloIT Systems ArchitectITCorporateFemaleLatino
E02006Harper DominguezDirectorEngineeringCorporateFemaleLatino
E02007Ezra VuNetwork AdministratorITManufacturingMaleAsian
E02008Jade HuSr. AnalystAccountingSpecialty ProductsFemaleAsian
E02009Miles ChangAnalyst IIFinanceCorporateMaleAsian
E02010Gianna HolmesSystem AdministratorITManufacturingFemaleCaucasian
E02011Jameson ThomasManagerITSpecialty ProductsMaleCaucasian
E02012Jameson PenaSystems AnalystITManufacturingMaleLatino
E02013Bella WuSr. AnalystFinanceSpecialty ProductsFemaleAsian
E02014Jose WongDirectorITCorporateMaleAsian
E02015Lucas RichardsonManagerMarketingCorporateMaleCaucasian
E02016Jacob MooreSr. ManagerMarketingCorporateMaleBlack
E02017Luna LuIT Systems ArchitectITCorporateFemaleAsian
E02018Bella TranVice PresidentEngineeringSpecialty ProductsFemaleAsian
E02023Lillian LewisTechnical ArchitectITResearch & DevelopmentFemaleBlack
E02024Serenity CaoAccount RepresentativeSalesManufacturingFemaleAsian
E02025Parker LaiVice PresidentAccountingSpecialty ProductsMaleAsian
E02026Charles SimmonsManagerSalesSpecialty ProductsMaleCaucasian
E02027Jayden LuuDirectorAccountingManufacturingMaleAsian
E02028Brooks RichardsonDirectorMarketingSpecialty ProductsMaleCaucasian
E02029Ivy ThompsonManagerMarketingManufacturingFemaleCaucasian
E02030Peyton WrightSr. ManagerMarketingCorporateFemaleBlack
E02031Wyatt DinhSystem AdministratorITSpecialty ProductsMaleAsian
E02032Ruby AlexanderVice PresidentFinanceResearch & DevelopmentFemaleCaucasian
E02033Axel OhSr. AnalystSalesCorporateMaleAsian
Encrypted Record
j7oYXhz4Wy5KpZV91BTBMQS3+73kDmFp6M8YCRVW4NXB1mQwCPSL13beBCSnIMkuyRtNHRzK05AO9kl7p745F3tbo7izTDb8NgVkCno+3vW5ET3uT8QoD27IBfClUEe7DNJn9dVL7aPOg1A3Rv6+YBmzOm+7KsW88IW7VWXdL/kMm5o4M8lbgAHrLO/xuWMEkloBl1lC1unF85heB7k8qXidalfK
40TVxVdOuPWk2evT097I/rl7vpqvCVaLRmx41Xd37arOTMDE/1pIGw+CF+TYuUkYOz7nRD+E8KgYkZOWNRQVtG9mI5a40gLV9ECCQufb2NoNJC4h387teruLlsm0S1vLKZNxH8jJeDlT3FahxXSsm3+dCu7lpaLsir323L5GBilEbmZmfq6yHvKUN1wl27v7XAheOzAsmWbAWEw9tPJqkR7ImdHR
ycfmVTVnwTd7JKbGhuKbui77502QGW2RyGwJvS+cbu/Z7HpAZAmiu3zbhOB9IaNEBdz7/GZf21hn/k47pWvYEKC/ivbPVSnFfujQDHQigjbaJipulksXPlNvaDHG7ju5IvLFZKf+vk2BhZuVz6AagupD7KgvPxmHnsDz7YbHM2d5zjUVTowWpxc7ZRUUswqY5l0Eg6DB4b8cYSaPKE2N0cNsTTqV
wPM0ULNFZTyEColRQOpvoHVxf9f9Sd2sdNY56f8c4cse+RSdZpxV5HeHauG4YktqUzdaAr+5N1ZUebGKNoWnszz1jxmBJWKDwBSDcrpejHVloy3+Q/haK1J9fhaN0VbQJ61UG67kSg4TLoDe8fwxUx70wCQvXh2+c60TyLPlyLrapA1YRWKiJ2DE7mc2ebOBqo1TNUyTFQViAl5rvXHE1F5Nd+GZ
sdGegbdXdAIgmPmODGp+nsws4yoOp6euic2nqqAS3VMCLu6WfhM8dmnsKEiBNW9mHj0H3SJXpFuE/YwVMS067PdCbIOn0+hxP/YwHL9jt2drGI3WvqaMMKH23em/u7RnxvzDgX1hhPnvfPLyJ8kwPyaUiShMgvuE6pYasVeHBoKaiBiOiuoM6EnA7mdKBGAcbZ+uz99SABScJMhvCN4lRk8ovZc9
s95HLdsYpbzZvioxBM5SmJy9IyVIHG3fxyOtaGn5LLH8QW3UH+/G7H2ErLbPmUbvB5/PX2/S35tTuGgHIWQ78nmNbsSj+4BFGgduKd9xCMvQ42UPPk3O0Z+Tx0LpL9NP80uAnjpPiOy0kDbw2BNi1TKXsWNaeIkMLG8mmk9Z+tgwNiapsJIbChOdwNxPqkOgWLLxDmE5RQJhH6hcmVVRgr2Fg3tB
eJhaXet+I1ofcvAe5bacN8kcD/CsTZckIWs8eerJNHcSJeHGBpDSK7SDJJTnEjKR8XVGgJ/0cZBEyzQmSpFVYYosa9VW00Lx0xTjzihPHoFBQBMGLMBsxKrxbhMZD3P5Og7hIHoVESwLZi0h7qGboGEkT4N+X3+JbORAtxczHWFJjNEtA4QqZptpVEw1iRWAPRjefJ7BuQ8krLoZhpGzaF8CPg09
LmQal/IFkKV9cw0mrynMkUeGU1R2/i9GseBDNQoCyuov6A8K1QE16ajNIADleXLeba3NjzE6fg8A61EzeRq0LGLH74zWePvVXm5uqEIX2lbTOi3qrhTIcXS06r/hn5NxmYpYQ5xwJAoQCSJhuZIs+UFvmT+1/p4Cc4PMpRpSyWRKeKOswR0UJKmqSt9T8DU9BGz2M7KPtdYddRIasPTzWILhoN8o
X+0DanOm8vzTsfda8qargbhLzP7B4gzdik0tgDqh0B57d/MD2MGpi81RACrJT3NLlWYTbtFwFYFR3NUoh8RVTryCQudlXsvj8s+P+5MMzbPtX9FJG3Ka7eZOG9FQmi5R9/l2QUOfirZLYannGTbD052lOmSQRj1vK+xHvCeeFamBfiVgmnigX9+hOnDO2dVROyOBsgG+hNF7kKtQKF/AXiNNRGGf
fbilPMzWe722qis0EdBJ0Kr+KNcO3KQXW4SGHSbqm+tN7F2UUxydDNDb3cy5xtI4Etv67zedtZ5dPZQaKLidJ+Z7EE/WZd3K61m9O7Iee4pS+VBM4fD5+d9QGxm1lr76MSEmU9mPme6u40bhb8TFHnyiv3otrWM3cYcS2p09yXjt0LwoHTM3VNEayPsk6f8wFFgMqG6rZAylDzp3mIlqZFcvAbVD
cWu6oCm7WlO1QjdzYR7ViYVD6BWmGTapda+48CKNQJ9fC+FmiIx7ODNHbddS/8zblbY5wYFMR8ozliTNtD5StjXISz/mgwcJQ78+KhJgN/E32Boc9wZYSamil026DElyIPBkTuvt3zjfnjjpkFxknUNAoLByXb/Wsja6nkmUNWwX0cIa/l/u0ZbLYVe+XhXlvJrkpnJIjE1aEvM9hKIcUe/VI2rl
UXYUKch7Zwf8u1WbJwLk8npTee0LBCaX9LUqyUKO3rDH/6dm3RKndomzXZXq57JH/ZhYWIhhauxAYxsHREPG/k4V/wG7SegnqVsahB2AnPery3Ju9WxZsWYWDzXs/kxFaJYpaYh3Q4VrR0AwV3BdGdYrGzo8OpZ0X6Ys9W4UUFiea2+AFroNEJpJ2yrL+zHDoC6loT814gQkdm0PWhEevn7JA5rQ
1KW52sH8l9rRGlDVezI1eXFDfCzlZjGNy3ZeoIb4mfAWOSWgVJdbrgcHP2uOHNf4MHs6tdvt2fQFYqgckxE78dmem7XZy93S+Eg/AluVm7geVUVN4LVBpBJ2gwLLhK21rUTCdutFUPIuNo1tT0DXfYC5dv/xAsEjC66m5EZRrGXtow+01DOC79rCQZOEMFJRcO3vngdGSh3LMlkvmjAcHJ7ZkptR
ukhwKa7c13QV5cxd0Ulcsv4it+noGlAXEY2PdAWidynP9WsMprR98+xajfQIG4fgl2z1UPNSup5edL65NrjAx5YbVsPO8GlB2m84sBKhkdlybGpYH8r16fJ0haZMZVl9ulf5/dYrSXOfNQN5UaVbg+564aiFRVx+X/Jl7bIc8HB4yhFCPeXud4jCbKZ/uGFO+iBhIxRS53T2nT3+WCJecXQyUZzb
18TLkAx7QnZEeZFQ65F0MmxsWa28tpjZuVYUQe4I/mEcm3198R+6URz4+oFBQbMo1IoWWgdWfAlNNBxcjKpw3AgezMjQkF6aYKZsu2b4V13tUYVwSrA7L1PKc9QzwUCgVWQV9jWLcufUBsbMjY6bmDIaSFXsHliLcyOP+IZ7YoK9jr0RO05w9XjsPrRlaNtAbX1zqcRtWktVfefQ6lbP4G1E+6gU
AEvcOlqBnTGthcVkRSrf159Rf2FpiJ9KGbmmw0CI/Ugk/nOs+YZaNOgMCV7pZmBK7mzE+Rjmaw2+H9xIFUScteCF0SMTvMweG0OuE7MAdzJYs8VsGGYUGutbYkKSpTOC7DNuAYrlnRmQ9oD0zkOipe2xauJfIl1OW9OYd3nQcjOHIEa9L2b6yOWPZLOejl9xioih/4Jb3sQoH5IcB8DoCkCkf9l5
hvdtCPSYVUUOLclO7E+Vpl8mEqzTpivnrNBJ4sUoRJghFP713JDecISUeCbPHMRGRC5cXe29wEr5vG2S4+WNOVANKT2FM8TJJ9ewuFMLP8JqofQFYFdajNHdRoIYAiJUGRr7TC93Rmq6IwrnI1zSnLrdzZccNuxD9vr6jA2ZPkPumK3n+AB63hBRmJKBeTiQdmeSOSMp1dxYOmZB80RgIHFobBAp
9Ns49OrjNoDjaDWA+7mMfdVyx86wh1RoDFdoS0Ih1+h0jHDtd1sMSEQVTSQGFmMmdMC8etGACqTjajwIji5YdWhJIwP4KDwG0JGEQJ6Tcsr9wwghjnU48brdj5UxXDqMPGI1PVPmiRNIkWuUWgX+B3qC9Wmb8v6mI2CN4XWmdOfJEVL6UbXm/OpXmskekYF62MskWBhJQmKY8dSTjo2c8/XA0mAV
qmDUliaXkTBLEvEdVTwO2wB2xzb6LfAwqN6jZ8ORJ3KWe4T9PFZ54sgeyhCxhrSzDDkomYXeESajRLOIC51fvD279ysazcVkurh4QXSziFyl2LlH0VjYEsZ7HwFDeoI/PhsVELOfKGMMoUdqxJCGmWBD2A8e4a4TVdHAQPqVn6e0mtVx2VMbbpO9o+CEqEAiLqq3726NSAmgCkRjQtXU8PrA5XWC
4EwKR5oBPJJxcrT2m/1IP/g+XjLYKJyKgmZscWuCZLokZYOjXTa4NKeheAKcbmcwHQ75lH0el6wGLQFl1p0Q2EpxKBPCuyxnloKFZ0pFKNrG+0NS9SrXv5Ab2S9RIdig/qgW9RXAj4YlFiOOHmDFk6mwgI8WeDOGL3JzKl+efNj7I5Yp4Xu7Fc8mZCq118EeMLiqIfxkgwq2uG7JKwQnuFWVlW6B
2LzpYSQY5gacZwORv0/CT4R7TaecpxiSNdCJh7vmw2HPj2nFPSuRizX7hFgASJr7AL4DDIMXiaZpdwOma/Yo0gvMA8N/LYQgMwGCXbGqWKufu+7qUmpybZJq8cK/NL8SbsABo4UDnVZmdEaXWfFPgVZhlER2wUah8/B4k959S97ObyODKFAjeQ4CA9zLPBeptWwO1Ub5x3o8jwUxNULyIc/4WxT9
nW7+ZQr6S+1bGOorj9BrJgzEWziJ4sXEVCVjwPsP82OahZHBlD2FqEBXjP1MdyENTO4fA0/dRpwBvmunCleFrYEZjGUkWXHk4p+JElZ/Ex9GZN1vZAqUbdDfpyePa9gLVbnX4acACNhApNCp1LB5Xpf0ZAwC+Ge2WcACAB8Hge72U1Ud5d6YlcLNiLrjN4mJOB0Mo9F7H2/CJMezR9zVzpchMq4d
Zmn+gC4x90YStRK+uYnqIhxe9TUOJ4OtFMCpU4LLMt87UK6NHSQSZgkjoXJ/23v188VAlZSPdwP74sS4vL7vinCBQM2nfTV4SlLOHyPdJb1/7e4kHj2bdSHcgV9e2CvV1vCtnQuhDM/TIdJCP/6lfzgxJ/PHqcicib4vR6A83/idEsWCD+x5U5gBq+bn6MTn/K07tpP3mX5uLpRNOCgg4x6xdm6Z
yaoJs+q7EpWRkrKRezkhpRjuVc+9972+umC/eoIo3QmAWLjyysyTJI6VLyQN5y7lZsQhBePcD9FdvkigFCvbnOZs1X0aHgFK8Cr2PUPOp/CgQhx+dseZwiKP6A/JDTr06THb+frWsGNy5ndgBbTqmGZwUGxQy1mf6b1p89Gw50Ley7i3J2IuUSnY9FSogvemDsPdcIuKULXikuh40uplQ+w+nycN
6GAyIQHExgrVZH6JvORg4Cknx69dI6PBLpKEwI3Usr87tL7C5A9cUqBJ5nVe6eUrkGYOMIe9IJ9FAp+iqj9QxLpu/kCEw599RoJJCPhyAGCnX4OL673zixyvrgmtUgLx9K3hYQmB+u/xDmU7tD/HP7mAdgdwMU/LLK3IW6nksqEA+QDR6F/POvA9mBsfTMJUfEu2eL3nkuZHLHr4WCphRer3lsWG
iQWsownWRVq9sIZVxv7OcdRUVEtdcJ7cpUd1aJ+/JPc/mRa99mg6kEFDZtRvJ9vHs5WVIBGDH1SKr1s0tddOp3Pimm4VK7VtTa8SYGHjnaNLsb9AX3BHZMoFd3DK3uxTodJneab2xns03Rkm3kWHyb9faPTHKU2lOgNPbsI3IcMj+bdh5TwKo3xxCsPBsPTg4AJrjUOywSl4WuF1VlXTe+jLpHyb
kPzUQiEs30VWytcSnGjzzpmVB6LRtuJIwnwUdTGv+ByinQnNC+Cd26nbtEH+iTZde4FukfbJrSkLa1zYTK/99xT3S+7OmP4gVJFEP8Zk4hIja3wcRrh8LtZnjvKEHP5Mgx0z8Or5K6ze35m+Nl3wv12bGonyCPY04eLZo5Vf48dXbTbJPIvAeaf4lvvBjSvGmWvZ2n4CtVl3jPk6QHBmVdajlGDK
bFyWhpsyETgL5mfVDWdMkoBYrcXTVWDBI6BzUAoKFJk3dwBBUJrNQAnF9GogwfxdWGOfTCbxElWuHG0XpsalJv7n6W/kYmoY4CmVLtriLTdveQiSYd1yUIRj+TfkQYrJA8HpXbAmtdNJ7reyZUjEkFMxerVDolnDmOYvonrhloHxryXWk2xA81NAg7KPKieIelZ1hRiRFBv0/mn3gzNU7vdecjdu
Raw unencrypted traditional database system
Encrypted Optimaize SearchCluster
← Drag the orange handle to compare ←→
Performance & Scalability

Built for billions of records, real-time response

10B
Person records
100B
Transaction records
10K/s
Updates & searches

Storage layer

OpenSearch (formerly Elasticsearch) provides built-in redundancy and unlimited scalability. Each data item is replicated across separate hardware, leveraging Apache Lucene underneath. Capacity scales far beyond global population figures.

Matching layer

Stateless matcher nodes scale horizontally across the cluster. They perform detailed record comparisons on candidates identified by OpenSearch queries. Since the layer is stateless, scaling is effectively infinite.

Typical deployment: 100M records, sub-100ms search response, 100 concurrent requests, writes visible within one second. Real-world numbers vary with data shape, query patterns, and infrastructure.

Redundancy & Backup

No single point of failure

Redundancy at every layer — data store, app nodes, hardware, data centers, and even API endpoints. Plus continuous off-site backup as a disaster-recovery safety net.

Data store

Records stored multiple times across separate hardware. Nodes can go offline without service interruption. Auto-sync, auto-scale. Over 10 years of operation without downtime using this technology.

App nodes

Stateless matcher nodes scale dynamically with demand. Nodes can be added or removed without affecting service availability.

Data centers & endpoints

Indices replicated across multiple data centers — surviving full DC outages from network failures, power loss, DDoS, or physical disasters. REST endpoints accessible through fully redundant URLs with different domains, IPs, and SSL certificates.

Continuous backup

Changes write through to an off-site incremental backup — safety net for human error, unauthorized access, software bugs, or large-scale hardware failure. Cross-cluster replication available for geo-distributed read access.

Record Versioning

Keep history when it matters

Customer records often need to preserve previous versions — name changes, address updates, identity transitions. SearchCluster gives you three patterns to choose from.

Default

Latest version only

Versioning disabled. Only the most recent write is indexed and returned. Applications can maintain history independently in a relational database or audit log.

Managed

Full record versions

New version stored on every update — incremental number + timestamp. Searches can target current, historical, or both. Retention strategies:

  • • Last N versions (e.g., 10)
  • • Time-based (e.g., past 3 years)
  • • Always keep the initial version
  • • Any combination of the above
Custom

Self-managed

Assemble records in JSON. Instead of storing a single email, maintain a list with timestamps: [{"email":..., "last-seen":...}].

Background Processing

Maintenance that doesn't hurt you

Predefined routines that traverse records without downtime, service interruption, or performance degradation.

Retention policy enforcement

Privacy regulations require deletion after defined inactivity periods. Bake compliance into the deployment — automated processes with continuous backup.

Encryption key rotation

The data encryption system uses background processing to rotate keys safely across the live record set.

Data model changes

When schemas expand with new fields, re-indexing may become necessary or beneficial — performed in the background.

Technical evolution

Search and match logic upgrades require internal re-indexing, scheduled in coordination with customers.

Record evolution

Each record carries an evolution version. On changes — key rotation, model updates — records are processed individually through separate transactions with optimistic locking. Processing happens distributed across cluster nodes. The procedure runs until all records reflect the new evolution number, potentially taking minutes to hours depending on dataset size. Crucially, the application works with both evolution versions during the transition — ensuring continuous operation throughout.

PII Compliance

Regulatory Requirements

Encryption

Record-level AES-256 encryption. No plain text data stored. Semantic search over encrypted data.

Audit + Traceability

Complete audit trail. Record versioning. Full traceability of all data access and modifications.

Continuity + Recovery

Continuous backups. Point-in-time recovery. Data lifecycle management and automated expiration.

European Union

GDPR - Data Protection
DORA - Digital Resilience
NIS2 - Cybersecurity
eIDAS - Digital Identity
AI Act - AI Compliance
Data Act - Data Sharing

United States

HIPAA - Healthcare Data
CCPA - California Privacy
GLBA - Financial Privacy
FERPA - Education Records
SOX - Corporate Data
PCI DSS - Payment Data
Challenges

Why Search Cluster?

Hard to Scale

Traditional databases struggle with billions of encrypted person records. Performance degrades with volume.

Slow queries at scale
Complex sharding required
High infrastructure costs

Hard to Prove Compliant

Meeting GDPR, DORA, NIS2 requires extensive auditing, encryption, and data lifecycle controls.

Manual audit trails
Incomplete encryption
Complex data retention

Hard to Guarantee Availability

Mission-critical systems need 99.99% uptime with automatic failover and disaster recovery.

Single points of failure
Manual backup processes
Slow recovery times
Technology

Semantic Search & Matching

Flexible Schema

person_id: string
name: {
first_name: encrypted_string
last_name: encrypted_string
}
email: encrypted_string
phone: encrypted_string
address: encrypted_object
metadata: flexible_json

Semantic Matching

Advanced AI-powered name matching finds duplicates even with typos, nicknames, and cultural variations.

Match Score: 94%
"John Smith" ≈ "Jon Smyth"

Search Over Encrypted Data

Query encrypted records without decryption. Proprietary indexing enables fast semantic search.

Query Time: 23ms
10M records searched
Performance

Enterprise Volumes

10B+
Person Records
Tested capacity
100B+
Transactions
Annual volume
10K
Ops/Second
Peak throughput
<50ms
Query Time
Average latency
Under the Hood

Technical deep dive

Open any topic to see how Search Cluster is engineered for scale, availability and operational resilience.

Performance & Scaling

Sub-50 ms search across hundreds of millions of records

Search latency. Complex search results against hundreds of millions of person records, with detailed match information, return in under 50 milliseconds. Typical 100M-record deployments respond within 100 ms while serving 100 concurrent requests.

Update propagation. Create, update and delete commands take effect throughout the distributed cluster within one second.

Scale. Production deployments handle 10 billion person records, 100 billion transaction records and 10k updates and searches per second. Storage runs on OpenSearch (Apache Lucene); the matching layer is a stateless application that scales horizontally across cluster nodes.

Exact numbers depend on data shape, search scenario and hardware setup.

Reliability & Disaster Recovery

Redundancy, off-site backup, cross-cluster replication

No single points of failure. Data is stored across multiple hardware nodes; some can go offline without service disruption. Stateless matcher nodes scale dynamically and can be added or removed without affecting availability. Optimaize has operated this infrastructure for over 10 years without downtime.

Multi-DC replication. Search Cluster indices are replicated across multiple data centres to maintain availability during full-facility outages — network failures, power loss, or other infrastructure issues. REST endpoints are reachable through separate, fully redundant URLs with different top- and second-level domains, IP addresses and SSL certificates.

Continuous backup. Changes to the data store write through to an off-site continuous incremental backup — protecting against accidental modification, unauthorised access, software bugs and large-scale hardware failure.

Cross-cluster replication. Maintain a read-only replica of your primary cluster for failover and geographically distributed query performance. Note: replication is not a substitute for backup — unwanted changes to the main cluster propagate to the replica.

Operations

Background processing & record versioning

Background maintenance. Long-running routines traverse records without service interruption: retention-policy enforcement (deleting customer data after a defined inactivity period to satisfy privacy law), encryption-key rotation, data-model migrations, and re-indexing for search/match logic upgrades.

Evolution versions. Each record carries an evolution-version field. During migrations, records are traversed and updated individually with optimistic locking; processing distributes across cluster machines and runs from minutes to hours depending on dataset size. Applications can operate against both versions during the transition.

Record versioning options.

  • Latest version only (default). The last write wins; previous versions are discarded.
  • Fully managed full record versioning. Each new write is stored with an incremental version number and timestamp. Search across current, historical or combined versions. Retention strategies: keep N recent versions, time-based windows (e.g. last 3 years), preserve initial version, or combinations.
  • Self-managed versioning. Embed your own version logic into the JSON record — e.g. an e-mail field stored as objects with last-seen timestamps.
Data Schema

Mapping your data into the cluster

1. Original data sources. Search Cluster accepts three shapes:

  • Tabular — flat rows and columns from SQL databases or CSV feeds.
  • Relational — multiple linked tables connected through keys and IDs.
  • Document — nested objects and arrays representing natural data structures.

2. Record feed. Records can be sent as simple CSV (spreadsheet-like rows), grouped CSV (using prefixes for entity namespaces), flat JSON (equivalent to CSV) or nested JSON (data from multiple joined tables).

3. Index schema. Internally the cluster organises data into entities with specific types (PERSON, ADDRESS) and fields with semantic data types (GIVENNAME, SURNAME, POSTALCODE).

Multiple values. Fields accept comma-separated strings or JSON arrays; structures support multiple instances via arrays. Combined with record versioning, the schema captures your data as it really is.

Integration

API-First Platform

REST API

Comprehensive RESTful API with JSON payloads. Full documentation and SDKs available.

GETPOSTPUTDELETE

Deployment Options

Cloud-hosted or on-premise. Your choice of infrastructure and data residency.

Optimaize Cloud
On-Premise Installation
Hybrid Setup

Example Integration

// Create encrypted person record
POST /api/v1/persons
Content-Type: application/json
Authorization: Bearer {token}
{
"name": {
"first_name": "John",
"last_name": "Smith"
},
"email": "john.smith@example.com"
}
Complete Capability Set

Everything you need, in one platform

Click any capability for details — or scan the grid for an overview.

Semantic record matching
Understanding person data makes all the difference. Match records across spellings, locales, transliterations, and partial information.
Record level encryption
No plain text data in any field. Every record is stored with strong AES-256 encryption — including indexes used for semantic matching.
Data maintenance
Record expiration, record evolution, lifecycle policies — the cluster manages retention and updates automatically.
Your schema
Simple flat or deeply nested data structure — any custom data structure works. Bring your existing model.
Your use case
Human-assisted processing with full-text search, and end-to-end automated processing — both supported on the same dataset.
Customize behavior
Input transformers, inferring data, match profiles — tune the matching engine to your domain without forking code.
Unlimited scaling
10 billion person records — check. 100 billion transaction records — check. 10k updates and searches per second — check.
Blazing fast
Complex search results with detailed match information in under 50 milliseconds. Sub-second response is the floor, not the ceiling.
Powerful APIs
Index, Update, Delete, Fulltext search, Typeahead, Copy search — and a versioned REST surface that won't break your integrations.
Fail-safe and redundant
Modern cluster technology with redundancy built in from the ground up — SLA with 99.99% uptime guarantee.
Continuous backup
Off-site backup for fast restore as disaster recovery. Point-in-time recovery for accidental writes or schema regressions.
Keeping history
Storing record versions — see how a record evolved over time, support audits, support right-to-be-forgotten.
Integration Options
REST API. On-premise (you run it) or SaaS (we run it). Run it where your data, compliance, and ops team prefer.
Ready when you are
Contact us to get started — kickoff usually within two weeks.
Why Now

The Time to Act is Today

Regulatory pressure is increasing. Data breaches are costly. Your customers expect protection. Search Cluster provides the infrastructure you need to stay compliant, secure, and scalable.

GDPR Fines
€20M+
Maximum penalties for non-compliance
Data Breaches
$4.88M
Average cost per incident
Detection Time
194 days
Industry average to detect breach