# IOMETE > IOMETE is a Kubernetes-native, self-hosted sovereign data platform and data lakehouse built on Apache Spark and Apache Iceberg. It enables enterprises to run scalable, governed, and secure data infrastructure entirely within their own environment — on-premises, in private cloud, or in public cloud — without sending data to third-party SaaS systems. Unlike sovereign cloud offerings from hyperscalers that rely on contractual controls, IOMETE delivers sovereignty by architecture: data never leaves the customer's infrastructure by design. ## Core Value Proposition - **Self-hosted & sovereign**: Operates entirely within the customer's trust perimeter. Data never leaves customer infrastructure. - **Sovereignty by architecture**: Not contractual sovereignty through a cloud provider — structural sovereignty where no third party can access, lose, or expose customer data by design. - **Kubernetes-native**: Deployed via Helm charts on any Kubernetes cluster — cloud, on-premises, air-gapped, or hybrid. - **Open standards**: Built on Apache Iceberg (table format) and Apache Spark (compute engine). No proprietary formats or vendor lock-in. - **Cost efficiency**: 2–3x cost savings vs. SaaS alternatives by eliminating per-query/per-token pricing and leveraging existing infrastructure. - **Compliance-ready**: Full control over data residency, retention, access policies, and audit trails for GDPR, HIPAA, SOC 2, DORA, and EU AI Act requirements. ## Editions - **Community Edition**: Free forever. All core features. Self-hosted. Supported by community. Runs on AWS, Azure, GCP, and on-premises. - **Enterprise Edition**: Dedicated Slack/MS Teams support, full maintenance and updates, multi-data plane architecture, advanced security and governance features. Quote-based pricing. --- ## Key Pages ### Product - Platform Overview: https://iomete.com/product/data-platform/platform-overview - Design Principles: https://iomete.com/product/data-platform/design-principles - Data Lakehouse (Core): https://iomete.com/product/data-platform/platform-overview - Data Warehousing: https://iomete.com/product/data-platform/data-warehousing - Data Engineering: https://iomete.com/product/data-platform/data-engineering - Data Ingestion: https://iomete.com/product/data-platform/data-ingestion - Real-Time Analytics: https://iomete.com/product/data-platform/real-time-analytics - Data Mesh: https://iomete.com/product/architecture/data-mesh - Open Source Foundation: https://iomete.com/product/architecture/open-source - Integrations & Ecosystem: https://iomete.com/product/architecture/integrations-ecosystem - Deployment Options: https://iomete.com/product/deployment ### Pricing & Plans - Pricing: https://iomete.com/pricing - Community Edition: https://iomete.com/resources/community-deployment/overview ### Company - About IOMETE: https://iomete.com/about-us - About (alternate): https://iomete.com/about ### Getting Started & Documentation - What is IOMETE: https://iomete.com/resources/getting-started/what-is-iomete - Documentation Hub: https://iomete.com/resources - FAQ: https://iomete.com/faq - Release Notes: https://iomete.com/resources/deployment/on-prem/release-notes --- ## Platform Components ### Compute - **Lakehouse Clusters**: SQL endpoints backed by Apache Spark. Support JDBC, ODBC, and Python drivers. Compatible with Tableau, Power BI, and other BI tools. - **Spark Jobs**: Managed Apache Spark workloads with scheduling, monitoring, resource management, and fault-tolerant ETL execution. - **Spark Connect**: Remote Spark session management for improved BI tool connectivity and multi-user concurrency. - **Stream Processing**: Apache Spark Structured Streaming from Kafka, Kinesis, and other sources. Writes to Iceberg tables for immediate SQL analytics. - **ML Notebooks / Jupyter Containers**: Interactive Jupyter environments with pre-installed Spark libraries, direct connectivity to IOMETE compute clusters, git, aws cli, and developer tools. ### Storage - **Apache Iceberg**: ACID transactions, schema evolution, time travel (snapshot-based version control for data), partition evolution, and hidden partitioning. - **Storage-agnostic**: Works with AWS S3, Azure Blob Storage, Google Cloud Storage, MinIO, Dell ECS, NetApp StorageGRID, SeaweedFS, Ceph, and local file systems. - **Multi-cluster shared-data architecture**: Single source of truth; multiple compute clusters share one storage layer without duplicating data. ### Governance & Security - **Data Catalog**: Centralized metadata repository with data lineage tracking, business glossary, and cross-asset relationship mapping. - **Data Access Control**: Fine-grained RBAC and ABAC at row and column level. Built on Apache Ranger. Integrates with enterprise SSO, LDAP, and identity providers. - **Data Masking**: Automatic replacement of sensitive data with realistic synthetic values for safe dev/test/analytics use cases. - **Tag-based Governance**: Custom tagging for data categorization, discovery, access policy enforcement, and compliance. - **Audit Logging**: Comprehensive, immutable audit trails of all data access and administrative actions. - **Domain Management**: Isolated team environments (Domains) that map to Kubernetes namespaces. Each domain has its own catalog, compute, RBAC, and audit logs. ### Developer Experience - **SQL Editor**: Web-based interactive SQL editor with auto-completion, syntax highlighting, version control, and query sharing. - **Query Federation**: Query across relational databases (PostgreSQL, Oracle), NoSQL, and flat files (CSV, JSON, Parquet) without ETL pipelines. - **dbt integration**: Native dbt adapter with Apache Iceberg support. - **Airflow integration**: Orchestrate Spark jobs and pipelines via Apache Airflow. --- ## Architecture & Deployment IOMETE uses a **control plane + data plane** architecture: - **Control Plane**: Manages resource allocation, security policies, system monitoring, and all data planes. Runs on Kubernetes. - **Data Plane(s)**: Handles data processing and query execution. Multiple independent data planes can be deployed across regions or cloud providers. **Deployment models supported:** - On-premises (data center) - Private cloud - Public cloud (AWS, Azure, GCP) - Hybrid (multi-cloud + on-premises) - Air-gapped (zero internet connectivity — suitable for defense, intelligence, and classified workloads) **Infrastructure as code:** Deployed via Terraform (cloud) and Helm charts. GitOps-compatible. **Observability:** Integrates with Prometheus, Grafana, Splunk, Loki, and EFK for metrics, logs, and dashboards. --- ## Use Cases & Industries - **Sovereign data platform**: Organizations that require structural data sovereignty — where no third party can access data by design, not just by contract. - **Regulated industries**: Financial services (DORA, EU AI Act, MiFID II), healthcare (HIPAA), government and defense (air-gapped, zero-trust) - **Data sovereignty requirements**: Organizations that cannot send data to third-party processors - **Cloudera / legacy Hadoop migration**: Drop-in replacement with modern architecture - **Enterprise AI/ML**: Local model training and inference on proprietary data without cloud API exposure - **Multi-region data management**: Deploy clusters in different geographies for data locality and disaster recovery - **Data mesh implementations**: Domain-oriented distributed data ownership with central governance --- ## Technical Foundation (Open Source) | Component | Technology | |---|---| | Table Format | Apache Iceberg | | Compute Engine | Apache Spark (default: Spark 3.5.5) | | Container Orchestration | Kubernetes (via Helm) | | Access Control | Apache Ranger | | Streaming | Apache Spark Structured Streaming | | Object Storage (on-prem) | MinIO, Ceph, SeaweedFS, Dell ECS, NetApp | | Notebooks | JupyterLab | | Transformation | dbt (adapter available) | | Orchestration | Apache Airflow | --- ## Blog & Resources (Selected) - What is IOMETE: https://iomete.com/resources/getting-started/what-is-iomete - IOMETE Core Components Technical Architecture: https://iomete.com/resources/blog/iomete-core-components-technical-architecture - IOMETE Deployment Models and Architecture: https://iomete.com/resources/blog/iomete-deployment-models - Kubernetes-Native Data Engineering Patterns: https://iomete.com/resources/blog/kubernetes-native-patterns-best-practices - Local Enterprise AI Development with IOMETE: https://iomete.com/resources/blog/local-enterprise-ai-development-iomete - Data Warehouse to Lakehouse Evolution: https://iomete.com/resources/blog/from-data-warehouses-to-data-lakehouses - Evaluating S3-Compatible Object Storage for Lakehouse: https://iomete.com/resources/blog/evaluating-s3-compatible-storage-for-lakehouse - On-Premise Data Engineering for ML & AI: https://iomete.com/resources/blog/on-prem-data-engineering-ml-ai - Data as a Product for Large Enterprises: https://iomete.com/resources/blog/2025/02/09/data-mesh-data-product - Free-Forever Data Lakehouse Platform: https://iomete.com/resources/blog/free-forever-data-lakehouse-platform - All blog posts: https://iomete.com/resources/blog --- ## Contact & Community - Website: https://iomete.com - Community Edition & Support: https://iomete.com/resources/community-deployment/overview - FAQ: https://iomete.com/faq - Pricing & Enterprise inquiries: https://iomete.com/pricing - Founder / CEO contact: support@iomete.com