Accelerate big data analytics by 100X or more.

We take the question to your data, not your data to the question.

Data has grown too heavy to haul.

Shipping it to a warehouse for every question clogs networks and burns compute.

Data lake Warehouse
Moving the data

So move the question to the data.

Our patented parallel processing selects, groups, aggregates and formats data in the lake, on industry-standard hardware.

Select Group Aggregate Format
Moving the question

Backed by federal research, built with partners.

Research funded by

Seven SBIR/STTR awards →

  • U.S. Department of Energy4 awards
  • NOAA2 awards
  • National Science Foundation1 award

University partners

Research partners, by field.

  • University of ChicagoAI image processing
  • University of Alabama in HuntsvilleMulti-dimensional data
  • University of Wisconsin–MadisonDatabase

Programs

Startup programs we belong to.

  • AWS ActivateSelected member
  • NVIDIA InceptionOn-premises solutions

Industry partners

Companies we work with.

Bring us your question.

Research

Federal research, done in the data lake.

SBIR and STTR awards fund our work on processing data where it lives.

7

SBIR and STTR awards

DOE 4 · NOAA 2 · NSF 1
5 Phase I · 2 advanced to Phase II

DOENOAANSF
Each dot is an award; filled dots are Phase II.

Five programs

  1. Scalable, Accelerated Genomic Sequence AnalysisDOE · 1 award
  2. H-WFQ for Isolation through Resource Allocation in HPC StorageDOE · 1 award
  3. Accelerated In-Storage Data Mining of Light SourcesDOE · 2 awards
  4. Accelerated in-storage analysis of multi-dimensional dataNOAA · 2 awards
  5. Real-Time Smart Data LakeNSF · 1 award

University partners

  • University of Chicago logoUniversity of ChicagoAI image processing
  • University of Alabama in Huntsville logoUniversity of Alabama in HuntsvilleMulti-dimensional data
  • University of Wisconsin–Madison logoUniversity of Wisconsin–MadisonDatabase

← Research

Scalable, Accelerated Genomic Sequence Analysis

Illustration
Problem
In genomic research for bioenergy, DNA sequencing now outpaces analysis: read mapping, aligning short sequences to a reference genome, has become the bottleneck.
Approach
A massively scalable, parallel read-mapping implementation inside AirMettle's analytic data storage platform, on standard commercial servers.

Award

Scalable, Accelerated Genomic Sequence Analysis

Agency
U.S. Department of Energy
Program
SBIR Phase I
Dates
February 18, 2025 – September 17, 2025
Amount
$256,500
Contract
DE-SC0026180
Official record on SBIR.gov ↗
Read the abstract

Genomic research is pivotal for advancing bioenergy applications, a key interest of the Department of Energy (DOE). Massive genomic datasets are essential for developing sustainable biofuels and understanding environmental impacts. However, a critical bottleneck exists in "read mapping"—the process of aligning short DNA sequences to a reference genome—where the speed of DNA sequencing far exceeds the capacity of data analysis. This mismatch delays bioenergy research advancements and hampers the potential of genomics to drive innovations in sustainable energy solutions. Building on proven, published research - the first in-storage processing system designed for genomic sequence analysis - this Phase I project aims to develop a massively scalable, parallel implementation using standard commercial server infrastructure. By leveraging advanced parallel processing capabilities in our existing analytic data storage platform, we intend to significantly accelerate the read mapping process. This innovative approach will enhance the efficiency of genomic data analysis, directly benefiting bioenergy research by enabling faster discovery and development of sustainable biofuels. The project will focus on three key areas: (1) Enhance the storage of large genomic datasets within the innovative analytic data storage platform to improve efficiency and speed; (2) Accelerate Read Mapping by developing solutions to perform genetic sequence matching while data is stored within a distributed storage environment; and (3) Validate how this solution can be utilized for bioenergy research. This work will be supported by the researcher who developed the in-storage genomic sequencing technology as well as a university supporting the bioenergy validation. If successful, this project could revolutionize genomic data processing, reducing the time and cost of sequencing, with broad applications in healthcare, agriculture, and bioenergy. These improvements would advance scientific research, potentially leading to significant economic and environmental benefits.

← Research

H-WFQ for Isolation through Resource Allocation in HPC Storage

Illustration
Problem
Multi-tenant HPC facilities run millions of processes at once, and existing resource-allocation methods struggle to share system resources dynamically and fairly under heavy load.
Approach
Integrate Hierarchical Weighted Fair Queueing (H-WFQ) into AirMettle's analytic data storage system to isolate tenants and processes with multi-level, dynamic, fair allocation.
Outcome
Released as the open-source HWF-Q Scheduler: MIT licensed, written in C11.HWF-Q Scheduler on GitHub ↗

Award

H-WFQ for Isolation through Resource Allocation in HPC Storage

Agency
U.S. Department of Energy
Program
SBIR Phase I
Dates
February 18, 2025 – September 17, 2026
Amount
$206,500
Contract
DE-SC0026122
Official record on SBIR.gov ↗
Read the abstract

High-performance computing (HPC) systems at the Department of Energy (DOE) and other research institutions are critical for large-scale scientific discovery. These inherently multi-tenant facilities, with millions of processes active simultaneously, necessitate robust mechanisms to ensure high performance and stringent security across multiple users. However, existing resource allocation methods often struggle to dynamically and fairly distribute system resources under heavy load, leading to performance bottlenecks and potential security vulnerabilities. To address these challenges, we propose integrating Hierarchical Weighted Fair Queueing (H-WFQ) into our existing analytic data storage system. By implementing H-WFQ, we will enhance user isolation by controlling and limiting resource access at both tenant and process levels, thereby significantly improving security and performance. During Phase I, we will develop methods to integrate H-WFQ support into our analytic storage system to strengthen user isolation. Specifically, we will design multi-level resource management mechanisms for dynamic, fair resource allocation among tenants and users. We will validate the enhanced system's effectiveness through performance and security testing in simulated HPC environments and ensure compatibility with the Globus data management services. This integration will significantly enhance the efficiency, security, and performance of HPC systems by providing robust isolation aligned with administrative policies. In Phase II, we aim to commercialize this solution for large multi-tenant HPC facilities. Strengthening user isolation protects sensitive data and ensures compliance with regulatory standards across industries, representing a breakthrough in resource management for HPC systems critical to scientific and technological advancements.

Open source · MIT · C11

HWF-Q Scheduler

Multi-level, work-conserving resource allocation and isolation for multi-tenant high-performance computing. Tenants and their processes get a fair, dynamically adjusted share of the system, even under heavy load.

  • Two-level hierarchical scheduling: tenants at the system level, flows within each tenant
  • WF2Q+ (Worst-Case Fair Weighted Fair Queueing) with precise fairness and bounded delay
  • Scales to 4,000+ tenants and millions of per-flow queues in under 2 GiB of memory
  • Work-conserving: unused capacity is redistributed; reconfiguration takes effect immediately
  • O(1) amortized enqueue and dequeue with a calendar queue of 16 exponential groups

Funded by DOE award DE-SC0026122

← Research

Accelerated In-Storage Data Mining of Light Sources

Illustration
Problem
X-ray and synchrotron light sources at DOE facilities generate datasets that reach petabyte scale, and their growth in volume and complexity makes them hard to process.
Approach
Mine the data where it is stored, with AirMettle's parallel in-storage processing, instead of moving it first.
Outcome
Initial HDF5 scientific-data support is in early-access trials in AirMettle Select.HDF5 in AirMettle Select (early access) ↗

Awards

Accelerated In-Storage Data Mining of Light Sources

Agency
U.S. Department of Energy
Program
SBIR Phase I
Dates
February 12, 2024 – September 11, 2024
Amount
$206,500
Contract
DE-SC0024750
Official record on SBIR.gov ↗
Read the abstract

High brightness X-ray/synchrotron light sources are employed in many experiments performed at Department of Energy (DOE) research facilities and in laboratories around the globe. Detectors and other telemetry produce Petabyte-scale datasets from these experiments stored as high-definition (HD) image files in Hierarchical Data Format Version 5 (HDF5). The complexity and production of experiment data is rapidly increasing at DOE's modern X-ray facilities, presenting enormous challenges in data acquisition, processing, analysis, storage, and management. These challenges, if left unresolved, will ultimately impact scientific productivity. AirMettle Inc. is developing a real-time smart data lake solution that simplifies big data analytics and accelerates processing by an order of magnitude, or more. AirMettle's key innovation involves massively parallel distributed data processing within an easily deployable software-defined storage framework ideal for modern cloud environments. The solution performs basic analytics tasks at the storage layer that reduce network traffic, improve data freshness, and enable real-time operation. This project will enable the storage service to efficiently ingest and then partition HDF5-formatted light source (HD image) data so it can be processed in a massively parallel manner in the storage layer. The teams will also invent the APIs necessary to enable AI-based analytics functions and models to be applied during the in-storage processing of this data. The resulting commercial software/service from this project will be a data storage solution that enables orders of magnitude faster ingestion and analysis of petabyte scale light source data sets, as well as image and video processing for public and private organizations. This should accelerate the pace of scientific research while enhancing industrial processes such as semiconductor manufacturing, leading to lower costs and higher quality consumer goods.

Accelerated In-Storage Data Mining of Light Sources

Agency
U.S. Department of Energy
Program
SBIR Phase II
Dates
April 14, 2025 – April 13, 2026
Amount
$1,150,000
Contract
DE-SC0024750
Official record on SBIR.gov ↗
Read the abstract

Modern X-ray and synchrotron light sources at Department of Energy (DOE) facilities generate vast datasets, often reaching petabyte scale. The growth in data volume and complexity presents significant challenges in processing, analyzing, storing, and managing these datasets efficiently. Addressing these challenges requires advanced tools that extract insights rapidly and securely while minimizing networking and storage bottlenecks. This project addresses these challenges by enhancing an analytical data platform to enable the rapid, in-place analytics on large image-rich datasets which are distributed across a large storage cluster on which the platform runs. In Phase I, we demonstrated the ability to partition large Hierarchical Data Format Version 5 (HDF5) datasets and to process them in parallel within this platform using User Defined Functions (UDFs), including AI inference models. We also validated the platform’s compatibility with tools commonly used by DOE researchers for accessing remote storage and computational resources. Phase II will expand the platform's capabilities to analyze a broader range of image-rich scientific and commercial datasets, improve its stability, security, and performance, and prepare the solution for commercial trials. The enhanced platform is expected to reduce analysis time for large image-rich datasets by up to 100x by leveraging massively parallel processing. This capability will empower researchers to derive actionable insights more quickly and efficiently, accelerating scientific discovery. Potential commercial applications include the analysis of drone surveillance data, with future opportunities in advanced manufacturing, medical imaging, and environmental monitoring. By addressing the growing demand for scalable analytics, the platform will foster innovation, enhance research efficiency, and support the development of data-intensive technologies.

← Research

Accelerated in-storage analysis of multi-dimensional data

Illustration
Problem
NOAA programs and the wider scientific community store massive volumes of multi-dimensional, array-oriented data, predominantly in the NetCDF standard.
Approach
Analyze that data in place, turning it into a real-time smart data lake for NOAA and the broader scientific community.
Outcome
NetCDF4 sub-selection and regridding for climate, ocean, atmospheric and geospatial data is in early-access trials in AirMettle Select.NetCDF4 in AirMettle Select (early access) ↗Multi-dimensional demonstration ↗

Awards

Accelerated in-situ analysis of multi-dimensional data

Agency
U.S. Department of Commerce, NOAA
Program
SBIR Phase I
Dates
September 1, 2022 – December 31, 2022
Amount
$150,000
Contract
NA22OAR0210591
Official record on SBIR.gov ↗
Read the abstract

The massive volumes of multi-dimensional array-oriented data generated by NOAA programs and the scientific community at large are predominantly stored in industry standard Network Common Data Form (NetCDF). Key challenges exist in making use of data stored in netCDF: data sets are often too large to be copied and transferred across networks for every user, and each time data is accessed by an analytics tool it must be retrieved, subsets extracted, and subsequently formatted, among other requirements, which can account for 80-90% of the total time needed to insight. To unlock the enormous potential of petabyte scale netCDF-formatted data stored at different locations, in this SBIR Phase I project, AirMettle, Inc. with its research partners from the University of Wisconsin-Madison proposes to explore the feasibility of integrating in-situ analysis capabilities for multi-dimensional data (netCDF) into our highly innovative real-time smart data lake solution. Dramatically accelerated data analytics performed at the storage layer addresses key challenges noted within the Climate Adaption and Mitigation Topic. Reducing data traffic between sites, shrinking required compute resources, and lowering costs – all while accelerating climate analyses by an order of magnitude – would bring great benefit to NOAA and the broader scientific community.

Accelerated in-storage analysis of multi-dimensional data

Agency
U.S. Department of Commerce, NOAA
Program
SBIR Phase II
Dates
August 1, 2023 – July 31, 2025
Amount
$650,000
Contract
NA23OAR0210342
Official record on SBIR.gov ↗
Read the abstract

AirMettle Inc. is transforming big data analytics for NOAA and the broader scientific community with a real-time smart data lake solution. Our innovative method utilizes massively parallel in-storage data processing within a versatile software defined storage framework, deployable on-premises or as a cloudbased service. Building upon our successful NOAA SBIR Phase I project, we strive to improve the handling of large, multi-dimensional NetCDF4 datasets vital to climate and weather forecasting. This project incorporates on-demand rescaling, allowing users to directly load only the required data at their desired resolution from the storage service. We will validate the benefits for climatologists and enhance the solution's commercial robustness. Our goal is to accelerate basic operations by 100x and reduce the data retrieved from storage by over 10x for typical requests. The potential commercial applications span meteorology and climatology across various sectors, making it easier and faster for experts to access and analyze crucial climate and weather data. AirMettle's cutting-edge data lake solution is poised to revolutionize big data analytics and garner widespread interest from both public and private stakeholders.

← Research

Real-Time Smart Data Lake

Illustration
Problem
Bottlenecks in existing data-analytics technologies keep the value of big data locked up for research, business and society.
Approach
A real-time smart data lake that processes data where it lives, to fundamentally accelerate the pace of data science.

Award

Real-Time Smart Data Lake

Agency
National Science Foundation
Program
STTR Phase I
Dates
February 15, 2022 – August 31, 2022
Amount
$256,000
Contract
2135007
Official record on SBIR.gov ↗
Read the abstract

The broader impact of this Small Business Technology Transfer (STTR) Phase I project will be to fundamentally accelerate the pace of data science. New solutions that alleviate or bypass the bottlenecks inherent in existing data analytics technologies are required to unlock the value contained in “big data” and bring greater benefits to research, business, and society at large. The potential benefits include improved quality control for manufactured goods, reduced fraud in financial transactions, and enhanced customization of consumer services. Many common software applications, such as search engines, online shopping, e-commerce, medical applications, and social networks, are backed by data analytical processing services. This project will enable massive data sets to be pre-processed directly by a shared storage service, then allow this capability to be efficiently utilized by client analytic applications. This project will accelerate transformation of analytical data to useful insights while reducing network congestion, simplifying complex analytics systems, and lowering information technology costs. This Small Business Technology Transfer (STTR) Phase I project examines the challenge of how to dramatically accelerate data-intensive computing problems by enabling large data objects to be processed directly in the storage layer then efficiently utilized by client applications. The technology will be built on a software defined data lake used for big data applications. Key technology will be added to enable JSON, one of the most widely used data interchange formats, to be processed in a distributed manner in-place within this storage solution with the objective of enabling client analytic applications to retrieve not the complete object but only the desired subset of content they require. The storage solution will be augmented to transform the data into a format that can be directly consumed by the clients. This will substantially increase efficiency within the storage itself and between analytics clients and the storage solution – while acceelrating data processing by bypassing bottlenecks. The result will be a smart data lake that can reduce network traffic, improve data freshness, and enable real-time operation – while accelerating big data analytics by an order of magnitude or more. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.

Company

The people behind AirMettle.

HeadquartersAirMettle, Inc., 2700 Post Oak Blvd., 21st Floor, Houston, TX 77056

Leadership

  • Photo of Donpaul Stephens
    Donpaul StephensFounder and CEO
    Bio

    Best known as the founder of Violin Memory, where he developed the original concept; hired and led the core team; raised venture capital from strategic, institutional, and individual investors; established key strategic partnerships in both supply chain and go-to-market; and developed multiple accounts.

  • Photo of Dan Orlow
    Dan OrlowBoard Member
  • Photo of Matt Youill
    Matt YouillFounder, Analytics
    Bio

    Best known as one of Betfair's founding engineers, where, as Chief Technologist, he built the platform that grew into one of the world's largest online gaming companies. Enabling that growth was an internally developed database that processed more transactions than the combined European stock exchanges and catalyzed the launch of Betfair's own financial exchange, LMAX.

Engineering

  • Photo of Josh Fuhs
    Josh FuhsFounding Engineer
  • Photo of Dean Sutherland
    Dean SutherlandDistinguished Engineer
  • Photo of Dhyakesh Soundararajan
    Dhyakesh SoundararajanSoftware Engineer
  • Photo of Ernest Villafana
    Ernest VillafanaSoftware Engineer
  • Photo of Hejie Huang
    Hejie HuangSoftware Engineer
  • Photo of Junjie Hua
    Junjie HuaSoftware Engineer
  • Photo of Mohit Anand
    Mohit AnandSoftware Engineer
  • Photo of Rupin Jairaj
    Rupin JairajSoftware Engineer
  • Photo of Sharthak Ghosh
    Sharthak GhoshSoftware Engineer
  • Photo of Craig Sanders
    Craig SandersSoftware Engineer
  • Photo of Hong Le
    Hong LeSoftware Engineer
  • Photo of Matthew Staab
    Matthew StaabSolution Architect

Sales & Marketing

  • Photo of Stefano Valdettaro
    Stefano ValdettaroGo-To-Market Strategist
  • Photo of Katsumi Kato
    Katsumi KatoSales Rep, Japan

Contact

Bring us your question.

Tell us where your data lives and what you want to ask of it.

Visit

AirMettle, Inc.
2700 Post Oak Blvd., 21st Floor
Houston, TX 77056

Customers

Support ↗