Skip to content

Commit f022fea

Browse files
Merge pull request #535 from arieljassan/feat/agent-runtime-generator
feat(evalbench): add native Agent Runtime generator support and deployment guide
2 parents ce34685 + b566d97 commit f022fea

10 files changed

Lines changed: 862 additions & 4 deletions

File tree

Lines changed: 41 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,41 @@
1+
############################################################
2+
### Dataset / Eval Items (Vertex AI Multi-Agent Test)
3+
############################################################
4+
dataset_config: datasets/bird/prompts.json
5+
6+
databases:
7+
- california_schools
8+
num_trials: 1
9+
10+
database_configs:
11+
- datasets/bat/db_configs/bigquery.yaml
12+
- datasets/bird/db_configs/sqlite.yaml
13+
dialects:
14+
- bigquery
15+
dialect: bigquery
16+
query_types:
17+
- dql
18+
dataset_format: bird-standard-format
19+
20+
############################################################
21+
### Prompt and Generation Modules
22+
############################################################
23+
model_config: datasets/model_configs/agent_runtime.yaml
24+
prompt_generator: 'NOOPGenerator'
25+
26+
############################################################
27+
### Scorer Related Configs
28+
############################################################
29+
scorers:
30+
python_scorer:
31+
script_path: 'evalbench/scorers/judges/hybrid_xa_judge.py'
32+
scorer_name: 'hybrid_cross_db'
33+
34+
############################################################
35+
### Reporting Related Configs
36+
############################################################
37+
reporting:
38+
bigquery:
39+
dataset_location: "US"
40+
gcp_project_id: !ENV ${EVAL_GCP_PROJECT_ID}
41+
csv: {}
Lines changed: 14 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,14 @@
1+
# Agent Runtime (Gemini Enterprise Agent Platform) Generator Manifest
2+
# Enables native NL2SQL evaluation benchmarking against deployed Agent Runtime endpoints.
3+
generator: agent_runtime
4+
5+
# Explicit resource URI string (e.g. "projects/12345/locations/us-central1/reasoningEngines/67890")
6+
# Leave empty string "" to dynamically read AGENT_ENGINE_RESOURCE from env.
7+
resource_name: !ENV ${AGENT_ENGINE_RESOURCE:""}
8+
9+
gcp_project_id: !ENV ${EVAL_GCP_PROJECT_ID}
10+
gcp_region: !ENV ${EVAL_GCP_PROJECT_REGION:us-central1}
11+
12+
# Rate limiting and retry mechanics
13+
execs_per_minute: 10
14+
max_attempts: 3

‎docs/configs/model-config.md‎

Lines changed: 26 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -28,6 +28,16 @@ These settings are **required only** for generators that utilize Google Cloud Ve
2828

2929
> Required*, you can globally set your GCP project_id and gcp_region using the environment variables `EVAL_GCP_PROJECT_ID` and `EVAL_GCP_PROJECT_REGION`.
3030
31+
## Agent Runtime (Gemini Enterprise Agent Platform) Configuration
32+
33+
These settings are **required only** for the `agent_runtime` generator, which connects to a live deployed Agent Runtime (Gemini Enterprise Agent Platform) instance.
34+
35+
| **Key** | **Required** | **Default Value** | **Description** |
36+
| ----------------- | ------------ | ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
37+
| `resource_name` | Yes | N/A | The fully-qualified resource ID of the deployed Agent Engine resource, e.g. `projects/<GCP_PROJECT>/locations/<REGION>/reasoningEngines/<RESOURCE_ID>`. |
38+
| `gcp_project_id` | Optional | `""` | The Google Cloud Project ID that hosts your Vertex AI resources. Can also be set via `EVAL_GCP_PROJECT_ID` environment variable. |
39+
| `gcp_region` | Optional | `""` | The Google Cloud region where the Vertex AI service is deployed. Can also be set via `EVAL_GCP_PROJECT_REGION` environment variable. |
40+
3141
## Important Notes
3242

3343
- **Customization:** This configuration is fully customizable to the needs of the selected generator. You can add or remove keys as necessary.
@@ -37,10 +47,11 @@ These settings are **required only** for generators that utilize Google Cloud Ve
3747
- **Rate Limiting & Retries:** The `execs_per_minute` and `max_attempts` keys help control the query generation process, ensuring that you can stay below project quota limits.
3848

3949

40-
## Example Configuration
50+
## Example Configurations
4151

42-
Below is an example of the updated YAML configuration file:
52+
Below are examples of the YAML configuration files:
4353

54+
### Gemini Model Example
4455
```yaml
4556
# General Generator Configuration
4657
generator: gcp_vertex_gemini
@@ -54,3 +65,16 @@ gcp_project_id: my_cool_gcp_project
5465
gcp_region: us-east5
5566
vertex_model: gemini-2.0-pro-exp-02-05
5667
```
68+
69+
### Agent Runtime Example
70+
```yaml
71+
# General Generator Configuration
72+
generator: agent_runtime
73+
execs_per_minute: 10
74+
max_attempts: 3
75+
76+
# Agent Runtime Configuration (Required for agent_runtime)
77+
resource_name: "projects/my-gcp-project/locations/us-central1/reasoningEngines/1234567890"
78+
gcp_project_id: "my-gcp-project"
79+
gcp_region: "us-central1"
80+
```

0 commit comments

Comments
 (0)