This file provides guidance for coding agents (including Claude Code and GitHub Copilot) when working with this repository.
Deephaven Community Core is a real-time, time-series, column-oriented analytics engine with relational database features. Queries operate seamlessly over both historical and live (ticking) data. The core engine is Java; it is consumed from Python, Groovy, and a set of client APIs (Java, Python, C++, JS, Go, R) that communicate over gRPC + Apache Arrow Flight + the Barrage streaming protocol.
This is a large Gradle multi-project monorepo (200+ subprojects). Java 11–25 is required; the build uses Gradle toolchain auto-provisioning to fetch the right JDK.
The build is driven by ./gradlew from the repo root. Subprojects are addressed by Gradle path
(e.g. :engine-table, :server-jetty-app, :extensions-parquet-table).
# Run the server (Python flavor) — serves the web IDE at https://localhost:10000/ide
./gradlew server-jetty-app:run
# Run the Groovy server instead
./gradlew server-jetty-app:run -Pgroovy
# Attach a debugger on port 5005 (combine with other flags)
./gradlew server-jetty-app:run -Pgroovy -Pdebug
# Build the Python wheel
./gradlew py-server:assembleThe PSK auth key is printed to the server log on startup; override with -Dauthentication.psk=<key>.
Most tests — including the engine suite — use JUnit 4, where @Category annotations route
tests to different Gradle test tasks. (Many newer client/extension modules use JUnit 5 (Jupiter)
via useJUnitPlatform().) The category-based routing below applies to the JUnit 4 tests.
There is no JUnit 3 left outside web/client-api, whose GWT suites require
junit.framework by construction. Do not extend junit.framework.TestCase, directly or through a
fixture: such a class runs under JUnit38ClassRunner, which finds methods by their test name
prefix and silently ignores @Test, @Before, @After and @Rule. Use @Test and, for the
engine fixture, extend RefreshingTableTestCase/QueryTableTestBase or declare
@Rule public final EngineCleanup base = new EngineCleanup(); when the class already has a
supertype. JMock comes from @Rule public final JMockRule jmock = new JMockRule();, and
io.deephaven.base.testing.Asserts supplies the array-aware and exact floating point
assertEquals overloads that org.junit.Assert lacks.
The default test task excludes the three categorized types (ParallelTest, SerialTest,
OutOfBandTest); each type has its own task (testParallel, testSerial, testOutOfBand) that
runs only that category. These categorized tasks are not wired into check — run them
explicitly by name. CI runs them nightly: the root nightly task and
.github/workflows/nightly-check-ci.yml invoke check, testParallel, testSerial, and
testOutOfBand directly. A pull request whose source branch is named engine/** or
coverage/** additionally runs testParallel, testSerial, and testOutOfBand on Java 21 as
part of .github/workflows/check-ci.yml, so auto-merge on such a PR is gated on them.
# Run the (uncategorized) tests for a module
./gradlew :engine-table:test
# Run a single test class or method (standard Gradle filtering)
./gradlew :engine-table:test --tests "io.deephaven.engine.table.impl.SomeTest"
./gradlew :engine-table:test --tests "*SomeTest.someMethod"
# Run the categorized test suites (invoke each task explicitly; CI runs them nightly)
./gradlew :engine-table:testParallel # @Category(ParallelTest.class)
./gradlew :engine-table:testSerial # @Category(SerialTest.class) — runs single-forked
./gradlew :engine-table:testOutOfBand # @Category(OutOfBandTest.class)
# Re-run tests even if cached/unchanged
./gradlew :module:test -PforceTest=true
# Faster engine tests with reduced data sizes
./gradlew :module:test -PshortTests=trueTest JVMs run with assertions enabled and use dh-tests.prop as the configuration root.
@Category(ParallelTest.class) opts a test into parallel execution.
Code style is enforced by Spotless (Google-derived style, see style/); generated checked-in
code is exempt.
./gradlew spotlessApply # auto-format (run before committing)
./gradlew spotlessCheck # verify formatting
./gradlew quick # lifecycle task: fast subset of check (compile + spotlessCheck)
./gradlew spotlessCheck quick # what the "quick CI" runs
./gradlew check # full verification (slow); CI runs with --continueWhen opening a PR, mirror the CI commands locally: ./gradlew spotlessCheck quick first, then
./gradlew check for the affected modules.
- Commit messages follow Conventional Commits, prefixed with the issue ID, e.g.
feat: DH-22670: Add codec-mapping abilityorfix: DH-22921: ...(enforced byconventional-pr-checkCI andcog.toml). - PR labels: every PR needs
ReleaseNotesNeeded/NoReleaseNotesNeededandDocumentationNeeded/NoDocumentationNeeded. - Each Gradle subproject must declare
io.deephaven.project.ProjectTypein itsgradle.properties. The type (e.g.JAVA_PUBLIC,JAVA_LOCAL,JAVA_APPLICATION,JAVA_PUBLIC_TESTING) selects a convention plugin frombuildSrc/that wires up publishing, testing, licensing, and dependency resolution. When creating a subproject, set this property and add it tosettings.gradle. .devin/rulesis the documentation style guide (for docs prose, not code).
The system is layered to decouple query syntax from execution, serialization, and client binding. Understanding these layers and how they connect is the key to navigating the code.
Table(io.deephaven.engine.table.Table, inengine/api, impl inengine/table): a columnar, typed, dynamically-updatable dataset.BaseTableis the abstract base for impls.ColumnSource<T>: per-column data accessor by row key. Tracks current and previous values (the "prev" mechanism) for change detection. Data is read in bulk asChunks (engine/chunk) — fixed-size, mostly zero-copy array windows — for efficient iteration.RowSet(engine/rowset): a compressed, possibly non-contiguous set of row keys.TrackingRowSetadditionally retains a previous-cycle snapshot.TableDefinition: the schema (column name → type).TableUpdate: emitted each cycle, describing added / removed / modifiedRowSets, row shifts (RowSetShiftData), and aModifiedColumnSetbitset of which columns changed.
Dependent tables form a DAG driven by the UpdateGraph. Tables are DynamicNodes;
they register TableUpdateListeners on upstream tables and receive TableUpdates via
onUpdate() each cycle.
- A
LogicalClocksequences cycle phases (Idle → Updating) so all listener callbacks within a cycle observe a consistent snapshot. NotificationQueueserializes listener execution;MergedListenercoordinates nodes with multiple parents.- Liveness / reference counting (
engine/liveness:LivenessNode,LivenessReferent,ReferenceCounted) keeps upstream dependencies alive through the weak-reference listener chain and ensures timely cleanup. NewTable-producing orListenercode typically must participate in liveness scoping. - Locking: operations annotated
@ConcurrentMethodbypass the update lock; others run under the graph's shared/exclusive lock within a cycle.UpdateGraph.serialTableOperationsSafe()gates thread-safety assumptions. - Attributes (
AttributeMap): semantic metadata propagated through operations (e.g.BlinkTable,AddOnly/AppendOnly,InputTablemarkers) used as hints/optimizations.
table-api(io.deephaven.api): provider-agnostic fluentTableOperationsinterface (filter, join, update, aggregations, sort, …) plus declarative types likeFilterandSelectable.qst(io.deephaven.qst, "query snapshot table"): an immutable query syntax tree.TableSpecrepresents a query graph;TableCreatorreplays a spec against a fluent backend. This makes queries serializable — the basis for remote/gRPC execution.engine/table/impl: concrete execution. Logical operations become graph nodes (SelectOrUpdateListener,SortListener, theByaggregation suite, join helpers likeCrossJoinHelper/AsOfJoinHelper, filter execution).
The engine processes large, ticking datasets on the hot path, so data-movement code must be written
for throughput. Before adding or changing engine internals (engine/table, engine/rowset,
engine/chunk, aggregation/join/update-by operators, ColumnSources, kernels), read
.github/instructions/query-engine.instructions.md — the rules cover bulk (chunked) reads,
dispatching to type-specialized kernels instead of per-cell virtual calls, allocating reusable
context objects before the per-chunk loop, batching RowSet operations, and keeping an operation
between two RowSets ideally O(n) while avoiding accidental quadratic paths.
serverexposes tables over gRPC + Arrow Flight; ticking data streams via the Barrage protocol (flatbuffer messages). Seeserver/src/.../arrowfor Flight handlers andBarrageMessageWriter.SessionState(io.deephaven.server.session) manages per-client object scopes and ticket resolution.- gRPC/protobuf definitions live in
proto/; Barrage flatbuffers and Flight bindings inextensions/barrageandextensions/flight-sql. py/serverbinds Python to the Java engine via jpy; Python authors queries that execute in-JVM.py/clientis the pure-Python gRPC client (pydeephaven).java-client/holds the Java client (session,flight,barrage) — note these use Dagger for dependency injection (the*-daggersubprojects), as does the server andplugin/dagger.- DI: the server and clients are wired with Dagger; look for
@Module/@Componentand the*-daggersubprojects when tracing how components are assembled. cpp-client/holds the C++ client:dhcore(client-side data model + Barrage ticking state machine, no Arrow/gRPC dependency) anddhclient(the user-facing API, over Arrow Flight + gRPC). It is also the substrate for two other clients —R/rdeephavenbindsdhclientthrough Rcpp andpy/client-tickingbindsdhcorethrough Cython — so changing a public header there can break them without breaking the C++ build.R/holds the R client (rdeephaven): R6 classes over an Rcpp module over the C++ client.
Before working in cpp-client/ or R/, read the corresponding design doc — each is written for
this purpose and will save you a lot of exploration:
cpp-client/DESIGN.md— architecture, code layout, the ticking pipeline, conventions, per-file summaries. (cpp-client/README.mdindexes the rest;cpp-client/BUILDING.mdis the build guide.)R/DESIGN.md— the R/Rcpp/C++ layering, the Arrow data path, conventions, per-file summaries. (R/README.mdindexes the rest;R/rdeephaven/BUILDING.mdis the build guide.)
Both use stable ## section anchors, so grep -n '^## ' <file> gives you a map to read selectively.
Server-side gRPC handlers turn untrusted client requests into engine operations, so they are a
security boundary. Before adding or changing a handler (server/src/.../table/ops/*GrpcImpl.java,
the hierarchical/partitioned/console/input-table services, or a service-loaded TicketResolver),
read .github/instructions/grpc-services.instructions.md — the checklist covers validating
every user-supplied expression through ColumnExpressionValidator, validating the exact string and
column shape the engine compiles, request-shape and authorization checks, error mapping, and the
tests to add.
extensions/: pluggable data integrations —parquet,kafka,csv,iceberg,s3,jdbc,json,arrow,barrage,protobuf,suanshu(math), etc.plugin/: server-side plugin system (object types, figures, hierarchical/partitioned tables).web/: the web IDE and JS client API (web-client-ui,web-client-api).Util,Configuration,IO,Base,log-factory: foundational utilities, theConfigurationproperty system, and the logging framework.engine/sqlandsql/: SQL front-end over the table engine.