GCO Regional API Gateway — API spec sheet¶
API Gateway api-gateway-regional · OpenAPI 3.0.3 · API version 1.0.0 · 10 endpoints · 0 schemas
Direct regional API for gco in us-east-1
One IAM-authenticated REST API per deployment region, fronting that region's internal ALB through VPC Lambdas. Its first job is to be the aggregator's path into the region: the resource policy always admits the cross-region aggregator's execution role for the aggregate routes. Direct calls by other principals of the deploying account are an explicit opt-in (api_gateway.regional_api_enabled) in the commercial aws partition and the required ingress in partitions without Global Accelerator. The routes are the same control-plane and inference catch-alls as the global API, with the additional HEAD and OPTIONS methods on /api/v1/{proxy+}.
- Machine-readable document:
docs/openapi/api-gateway-regional.json— the documentscripts/generate_api_gateway_openapi.pyreads out of the synthesized CDK stack; this sheet is rendered from it - Interactive console (Swagger UI): https://aws-solutions-library-samples.github.io/global-capacity-orchestrator-on-aws/swagger/api-gateway-regional/
- Catalogue index: README.md · interaction diagram
Servers¶
https://{api_id}.execute-api.{region}.{url_suffix}/{stage}
Stage prod of the deployed REST API. The full URL is the RegionalApiEndpoint output of the gco-regional-api-us-east-1 stack.
| Variable | Default | Description |
|---|---|---|
api_id |
"<api-id>" |
REST API id assigned at deploy time. |
region |
"us-east-1" |
One API per entry of deployment_regions.regional in cdk.json. |
stage |
"prod" |
— |
url_suffix |
"amazonaws.com" |
The partition's DNS suffix (${AWS::URLSuffix}). |
Deployment¶
- Endpoint type:
REGIONAL - Stage:
prod— throttling 1000 req/s (burst 2000), execution loggingINFO, data trace off, metrics on, X-Ray tracing on, access logs on
Resource policy¶
Statements deployed on the REST API (${AWS::AccountId} is the deploying account):
[
{
"Action": "execute-api:Invoke",
"Effect": "Allow",
"Principal": {
"AWS": "arn:aws:iam::${AWS::AccountId}:role/gco-cross-region-aggregator"
},
"Resource": [
"execute-api:/*/DELETE/api/v1/jobs",
"execute-api:/*/GET/api/v1/health",
"execute-api:/*/GET/api/v1/jobs",
"execute-api:/*/GET/api/v1/policy",
"execute-api:/*/GET/api/v1/status"
]
}
]
Added when api_gateway.regional_api_enabled is true (and always outside the aws partition), admitting the deploying account's other principals:
[
{
"Action": "execute-api:Invoke",
"Condition": {
"ArnNotEquals": {
"aws:PrincipalArn": "arn:aws:iam::${AWS::AccountId}:role/gco-cross-region-aggregator"
},
"StringEquals": {
"aws:PrincipalAccount": "${AWS::AccountId}"
}
},
"Effect": "Allow",
"Principal": {
"AWS": "*"
},
"Resource": "execute-api:/*"
}
]
Security schemes¶
sigv4¶
- Type: apiKey — header
Authorization(awsSigv4)
AWS Signature Version 4 for service execute-api, signed with IAM credentials of the deploying account that carry execute-api:Invoke on this API (for example awscurl --service execute-api --region <region>).
Endpoints¶
| Method | Path | Summary | Tags |
|---|---|---|---|
GET |
/api/v1/{proxy+} |
Read a control-plane resource | control-plane |
POST |
/api/v1/{proxy+} |
Create a control-plane resource or submit a job | control-plane |
PUT |
/api/v1/{proxy+} |
Replace a control-plane resource | control-plane |
PATCH |
/api/v1/{proxy+} |
Update a control-plane resource | control-plane |
DELETE |
/api/v1/{proxy+} |
Delete a control-plane resource | control-plane |
HEAD |
/api/v1/{proxy+} |
Probe a control-plane resource | control-plane |
OPTIONS |
/api/v1/{proxy+} |
Describe the methods a control-plane path accepts | control-plane |
GET |
/inference/{proxy+} |
Read an inference endpoint's health, info or model list | inference |
POST |
/inference/{proxy+} |
Generate: forward a completion request and stream the response | inference |
HEAD |
/inference/{proxy+} |
Probe an inference endpoint | inference |
Endpoint details¶
GET /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
- Operation ID:
getApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
regional-api-proxy(lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM /
/alb-hostname- , verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s. - Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
POST /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
- Operation ID:
postApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
regional-api-proxy(lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM /
/alb-hostname- , verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s. - Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
PUT /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
- Operation ID:
putApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
regional-api-proxy(lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM /
/alb-hostname- , verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s. - Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
PATCH /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
- Operation ID:
patchApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
regional-api-proxy(lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM /
/alb-hostname- , verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s. - Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
DELETE /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
- Operation ID:
deleteApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
regional-api-proxy(lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM /
/alb-hostname- , verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s. - Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
HEAD /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
- Operation ID:
headApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
regional-api-proxy(lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM /
/alb-hostname- , verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s. - Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
OPTIONS /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
- Operation ID:
optionsApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
regional-api-proxy(lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM /
/alb-hostname- , verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s. - Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
GET /inference/{proxy+}¶
Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.
- Operation ID:
getInferenceProxy - Tags: inference
- Security:
sigv4 - Integration:
aws_proxy, 900 s integration timeout, response streaming
Backend
- Lambda:
inference-streaming-proxy(lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout) - Behaviour: Node.js response-streaming VPC Lambda (ROUTING_MODE=regional): the same handler as the global route, reaching the ALB directly instead of through Global Accelerator.
- Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist —
inference-proxy - The endpoint's model-server Service in gco-inference (
, -canary or -proxy), streamed back through every hop
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
404 |
Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
502 |
Bad gateway: the in-cluster proxy or the model server did not answer. | — |
POST /inference/{proxy+}¶
Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.
- Operation ID:
postInferenceProxy - Tags: inference
- Security:
sigv4 - Integration:
aws_proxy, 900 s integration timeout, response streaming
Backend
- Lambda:
inference-streaming-proxy(lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout) - Behaviour: Node.js response-streaming VPC Lambda (ROUTING_MODE=regional): the same handler as the global route, reaching the ALB directly instead of through Global Accelerator.
- Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist —
inference-proxy - The endpoint's model-server Service in gco-inference (
, -canary or -proxy), streamed back through every hop
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
404 |
Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
502 |
Bad gateway: the in-cluster proxy or the model server did not answer. | — |
HEAD /inference/{proxy+}¶
Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.
- Operation ID:
headInferenceProxy - Tags: inference
- Security:
sigv4 - Integration:
aws_proxy, 900 s integration timeout, response streaming
Backend
- Lambda:
inference-streaming-proxy(lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout) - Behaviour: Node.js response-streaming VPC Lambda (ROUTING_MODE=regional): the same handler as the global route, reaching the ALB directly instead of through Global Accelerator.
- Then:
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist —
inference-proxy - The endpoint's model-server Service in gco-inference (
, -canary or -proxy), streamed back through every hop
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
404 |
Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
502 |
Bad gateway: the in-cluster proxy or the model server did not answer. | — |
Schemas¶
This api gateway declares no component schemas.