Skip to content

Bug: InsertStream throughput collapses 20-30x after a few hundred unary executeQuery calls #4197

Description

@TheRealMcoy

Description

An InsertStream call that takes ~0.3s for 100k rows in isolation reproducibly takes ~8.7s for the same workload if a few hundred unary ExecuteQuery RPCs ran first against the same server. The slowdown:

  • Persists across fresh gRPC channels — closing the channel that ran the unary queries and opening a brand-new one before InsertStream does not restore throughput. The bad state is server-side, not client-side.
  • Reproduces immediately after as few as 500 unary ExecuteQuery calls; once "tainted," every subsequent InsertStream on that server stays slow until the server process restarts.
  • Is unrelated to dataset content. The unary queries can target any vertex/document type; the bulk insert target can be a freshly-DDL'd type with no rows.

Environment

  • ArcadeDB 26.5.1-SNAPSHOT (build 15f105414cce73b90033c14c925e7f181a11aa3e)
  • gRPC InsertStream with InsertOptions.transaction_mode = PER_STREAM, server_batch_size = 1000
  • Default JVM settings

Likely root cause

ArcadeDbGrpcService.executeQuery is the only ResultSet-acquiring handler in ArcadeDbGrpcService.java that does not use try-with-resources:

Handler Line ResultSet lifecycle
executeCommand 316 try (ResultSet rs = db.command(...))
executeQuery 881 ResultSet resultSet = database.query(...); — never closed
streamCursor 1267 try (ResultSet rs = db.query(...))
streamMaterialized 1335 try (ResultSet rs = db.query(...))
streamPaged 1402 try (ResultSet rs = db.query(...))
(upsert lookup) 2137 try (ResultSet rs = ctx.db.query(...))

So every unary executeQuery leaks the ResultSet (and whatever page-cache references / cursor state it transitively holds). The leaked state accumulates per call against the same Database instance and gradually starves the path that InsertStream exercises (transactional writes via InsertContext.flushCommit).

The hypothesis matches every observation we have:

  • The leak is on the Database object itself, which is pooled in databasePool (line 142) and shared across channels — explains why a fresh gRPC channel doesn't reset throughput.
  • The leak grows with unary-query count, not unary-query content — explains the linear-in-N degradation we see (500 calls is enough to make 100k-row InsertStream take 8s vs 0.3s).
  • The fix is one line: wrap the ResultSet in executeQuery in a try (ResultSet resultSet = database.query(...)) { ... } like every other handler in the file already does.

Suggested fix

Wrap the ResultSet in executeQuery (ArcadeDbGrpcService.java:881) with try-with-resources, matching every other handler in the file:

// Before (line 881):
ResultSet resultSet = database.query(language, request.getQuery(), queryParams);
// ... iterate ...
// (method ends; resultSet never closed)

// After:
try (ResultSet resultSet = database.query(language, request.getQuery(), queryParams)) {
    // ... iterate ...
}

This should be enough on its own; if not, the same audit should be repeated against any other Database.query(...) / Database.command(...) call site in grpcw/ that doesn't close its ResultSet.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions