Streaming Orders and History into Apache Iceberg
Overview
Apache Iceberg is an open table format for large analytical datasets. Ember can stream trading messages and completed orders into Iceberg tables stored on Amazon S3 (or any S3-compatible storage). Data files are written as Parquet directly by Ember, and the tables can be queried by Athena, Trino, Spark, Flink, DuckDB and other engines that understand Iceberg.
Compared to the plain Amazon S3 / Athena warehouse, Iceberg gives you:
- a real table with a typed schema, no
CREATE EXTERNAL TABLEandMSCK REPAIR TABLE; - atomic commits: readers never see a partially written batch;
- partition pruning by journal term and day, and file-level min/max statistics for fast lookups;
- table maintenance with standard Iceberg tools, that are automatic with AWS Glue catalog.
The Iceberg warehouse is available only in Ember 1.15.28 and later, and requires Java 21 or newer to run the Ember data warehouse service.
Prerequisites
- An S3 bucket (or an S3-compatible endpoint such as MinIO) for the table data.
- An Iceberg catalog, one of:
- JDBC catalog (default): any JDBC database, for example PostgreSQL. Ember creates the catalog tables on first use.
- AWS Glue catalog: no extra infrastructure, and the tables are immediately visible in Athena.
- Permissions to read, write and delete objects in the bucket. For Glue, also permissions to create and update the database and table. See Recommended IAM Permissions for the S3 part.
Configure the Ember
Add the following section to $EMBER_HOME/ember.conf. The iceberg warehouse is defined in ember-default.conf; you only override what you need.
warehouse {
iceberg {
live = true
commitPeriod = 10m # see "Commit Period" below
messages.loader.settings {
warehouse = "s3://ember-warehouse/iceberg"
region = "us-east-2"
catalogImpl = "org.apache.iceberg.aws.glue.GlueCatalog"
maxBatchSize = 64K
}
orders.loader.settings {
warehouse = "s3://ember-warehouse/iceberg"
region = "us-east-2"
catalogImpl = "org.apache.iceberg.aws.glue.GlueCatalog"
maxBatchSize = 64K
}
}
}
Example with a PostgreSQL JDBC catalog:
messages.loader.settings {
warehouse = "s3://ember-warehouse/iceberg"
region = "us-east-2"
uri = "jdbc:postgresql://db-host:5432/iceberg"
jdbcUser = "ember"
jdbcPassword = "******"
accessKey = "******"
secretKey = "******"
}
Run the service the same way as any other warehouse, passing the unit id:
/deltix/ember/bin/data-warehouse iceberg
Loader settings
The same settings apply to the messages and orders loaders.
| Setting | Default | Description |
|---|---|---|
warehouse | - | Warehouse root location, e.g. s3://bucket/path. Required. |
catalogName | ember | Iceberg catalog name. |
namespace | ember | Table namespace (database). Use dots for nested namespaces. |
tableName | messages / orders | Table name. |
catalogImpl | org.apache.iceberg.jdbc.JdbcCatalog | Catalog implementation. Use org.apache.iceberg.aws.glue.GlueCatalog for AWS Glue. |
ioImpl | org.apache.iceberg.aws.s3.S3FileIO | File IO implementation. |
uri | - | Catalog URI. The JDBC URL for the JDBC catalog; not needed for Glue. |
jdbcUser, jdbcPassword | - | JDBC catalog credentials. |
region | - | AWS region for S3 access. |
s3Endpoint | - | Custom S3 endpoint (MinIO, LocalStack). |
s3PathStyleAccess | false | Path-style S3 access, required by most S3-compatible endpoints. |
accessKey, secretKey | - | S3 API credentials for data file access. Not used for Glue API calls, see below. |
properties | - | Extra catalog properties as key1=value1;key2=value2. |
compression | GZIP | Parquet compression: GZIP or UNCOMPRESSED. |
minBatchSize | 0 | Minimum number of records for a soft commit. A hard commit ignores it. |
maxBatchSize | 64K | Maximum number of records per data file. Must be greater than zero. |
For production setups, store secrets in the Hashicorp Vault or hash them using the Mangle tool. For more information, refer to the Ember Configuration Guide.
Credentials and AWS Glue
With GlueCatalog, the Glue API is always accessed with the AWS default credential chain (environment variables, shared credentials file, or the EC2 instance profile); accessKey and secretKey do not apply to it. They are passed only to the S3 file access. The recommended setup is to leave both unset, so that Glue and S3 both use the default chain; see Using AWS instance profile instead of API keys.
Commit Period
Every commit is a catalog round trip and creates a new Iceberg snapshot with new metadata files. The default commitPeriod of the warehouse unit is 5 seconds, which is too frequent for Iceberg: it produces many tiny data files and a long snapshot history, and degrades query planning.
Set commitPeriod to 5 to 15 minutes. Data becomes visible to queries after the next commit, so expect a delay of up to that period. Data files are also committed when they reach maxBatchSize records, and when the service stops.
Tables
Ember creates the namespace and the tables on the first run. An existing table is never modified.
| Table | Contents |
|---|---|
messages | All trading messages (requests and events), the same as the S3 messages data. |
orders | The final state of each closed order. |
Column names match the JSON format described in Appendix A: Data Format, for example Term, Sequence, Timestamp, OrderId, Symbol, LimitPrice. Custom order attributes are stored in an Attributes list of (Key int, Value string) structs.
- Partitioning: by
Term(the journal term) and by day of the record time (UTC).messagesusesTimestamp;ordersusesCloseTime. - Sorting: data files are sorted by the time column.
- Format: Parquet, compressed with GZIP by default.
Data Types
| Ember type | Iceberg type |
|---|---|
| Prices, quantities, commissions | decimal(38,12) |
| Timestamps | timestamptz (microsecond precision) |
| Term, sequence | long |
| Enumerations, symbols, identifiers | string |
| Flags and reject codes | int, boolean |
| Absent values | Parquet null |
Absent integer values are stored as real nulls, not as the -2147483648 sentinel used by the JSON format.
Querying
Athena, with the Glue catalog:
SELECT Timestamp, OrderId, Symbol, Side, TradePrice, TradeQuantity
FROM ember.messages
WHERE Term = 1672826718396
AND Timestamp >= TIMESTAMP '2026-09-01 00:00:00 UTC'
AND Timestamp < TIMESTAMP '2026-09-02 00:00:00 UTC'
AND Type = 'OrderTradeReportEvent';
Filter on Term and on the time column to benefit from partition pruning. You can identify any record by the pair (Term, Sequence), see Message Identity.
Behavior and Guarantees
Day split
A data file never spans two UTC days. When a record of the next day arrives, the current file is committed and a new one is started, so each file belongs to exactly one partition.
Restart and recovery
Each commit records the last exported Term and Sequence in the snapshot summary (keys ember.term and ember.last-sequence). On start, the service reads them from the latest snapshot and continues from the next journal record, so no data is lost or duplicated on restart.
If the latest snapshot does not carry these keys (for example, it was produced by a table maintenance job), the service looks at earlier snapshots and then at the statistics of the data files. If the restart point still cannot be determined, the service stops with an error instead of guessing, so you can decide how to proceed. Nothing is skipped or duplicated silently.
A table rollback combined with snapshot expiry can make the recorded restart point misleading. Avoid expiring snapshots that the loader may still need, and do not modify the table manually while the warehouse service is running.
Importing data back into the journal
The Journal Importer tool can read Iceberg tables. Records are replayed in journal order.
Table Maintenance
Over time a table accumulates many small files and old snapshots. With AWS Glue you can enable the built-in table optimizers (compaction, snapshot retention and orphan file cleanup) and they run automatically, with no jobs for you to schedule. This is the recommended setup.
With other catalogs, run standard Iceberg maintenance periodically (compaction and snapshot expiry) using Spark, Trino or Athena.
Maintenance is safe for Ember. It may rewrite files and snapshots, and the service still finds its restart point.
Limitations
- Decimal range and scale: decimals are
decimal(38,12). Values are rounded to 12 decimal places, and values beyond ±1e26 are clamped to the largest representable value, so the export never fails because of a number. - Timestamp precision: nanosecond timestamps are truncated to microseconds.
- Ember 1.15.28+ and Java 21+ are required.
Iceberg or S3 JSON?
| S3 JSON | Iceberg | |
|---|---|---|
| Table setup | Create the Athena table yourself | Created automatically |
| New days of data | Partitions must be added regularly, and Athena limits the number of partitions per table | Nothing to do |
| Loading speed | Faster, higher sustained throughput | Slower, data is committed in periodic batches |
| Data visibility | Right after each upload | After each commit, within minutes |
| Half-written data | Possible to see partial batches | Never visible |
| Query speed and cost | Scans whole files | Reads only the needed data |
| Upkeep | Manual | Automatic with Glue |
| Extra infrastructure | None | Glue or a JDBC database |
| Ember version | Any | 1.15.28+ |
| Java version | 17+ | 21+ |
Choose S3 JSON for the highest ingest rate and the simplest setup. Choose Iceberg when you want a maintenance-free analytical table. For Parquet files without Iceberg, see Parquet Format.