Skip to content

GCO Regional API Gateway — API spec sheet

API Gateway api-gateway-regional · OpenAPI 3.0.3 · API version 1.0.0 · 10 endpoints · 0 schemas

Direct regional API for gco in us-east-1

One IAM-authenticated REST API per deployment region, fronting that region's internal ALB through VPC Lambdas. Its first job is to be the aggregator's path into the region: the resource policy always admits the cross-region aggregator's execution role for the aggregate routes. Direct calls by other principals of the deploying account are an explicit opt-in (api_gateway.regional_api_enabled) in the commercial aws partition and the required ingress in partitions without Global Accelerator. The routes are the same control-plane and inference catch-alls as the global API, with the additional HEAD and OPTIONS methods on /api/v1/{proxy+}.

Servers

  • https://{api_id}.execute-api.{region}.{url_suffix}/{stage}

Stage prod of the deployed REST API. The full URL is the RegionalApiEndpoint output of the gco-regional-api-us-east-1 stack.

Variable Default Description
api_id "<api-id>" REST API id assigned at deploy time.
region "us-east-1" One API per entry of deployment_regions.regional in cdk.json.
stage "prod" —
url_suffix "amazonaws.com" The partition's DNS suffix (${AWS::URLSuffix}).

Deployment

  • Endpoint type: REGIONAL
  • Stage: prod — throttling 1000 req/s (burst 2000), execution logging INFO, data trace off, metrics on, X-Ray tracing on, access logs on

Resource policy

Statements deployed on the REST API (${AWS::AccountId} is the deploying account):

[
  {
    "Action": "execute-api:Invoke",
    "Effect": "Allow",
    "Principal": {
      "AWS": "arn:aws:iam::${AWS::AccountId}:role/gco-cross-region-aggregator"
    },
    "Resource": [
      "execute-api:/*/DELETE/api/v1/jobs",
      "execute-api:/*/GET/api/v1/health",
      "execute-api:/*/GET/api/v1/jobs",
      "execute-api:/*/GET/api/v1/policy",
      "execute-api:/*/GET/api/v1/status"
    ]
  }
]

Added when api_gateway.regional_api_enabled is true (and always outside the aws partition), admitting the deploying account's other principals:

[
  {
    "Action": "execute-api:Invoke",
    "Condition": {
      "ArnNotEquals": {
        "aws:PrincipalArn": "arn:aws:iam::${AWS::AccountId}:role/gco-cross-region-aggregator"
      },
      "StringEquals": {
        "aws:PrincipalAccount": "${AWS::AccountId}"
      }
    },
    "Effect": "Allow",
    "Principal": {
      "AWS": "*"
    },
    "Resource": "execute-api:/*"
  }
]

Security schemes

sigv4

  • Type: apiKey — header Authorization (awsSigv4)

AWS Signature Version 4 for service execute-api, signed with IAM credentials of the deploying account that carry execute-api:Invoke on this API (for example awscurl --service execute-api --region <region>).

Endpoints

Method Path Summary Tags
GET /api/v1/{proxy+} Read a control-plane resource control-plane
POST /api/v1/{proxy+} Create a control-plane resource or submit a job control-plane
PUT /api/v1/{proxy+} Replace a control-plane resource control-plane
PATCH /api/v1/{proxy+} Update a control-plane resource control-plane
DELETE /api/v1/{proxy+} Delete a control-plane resource control-plane
HEAD /api/v1/{proxy+} Probe a control-plane resource control-plane
OPTIONS /api/v1/{proxy+} Describe the methods a control-plane path accepts control-plane
GET /inference/{proxy+} Read an inference endpoint's health, info or model list inference
POST /inference/{proxy+} Generate: forward a completion request and stream the response inference
HEAD /inference/{proxy+} Probe an inference endpoint inference

Endpoint details

GET /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

  • Operation ID: getApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: regional-api-proxy (lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM //alb-hostname-, verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

POST /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

  • Operation ID: postApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: regional-api-proxy (lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM //alb-hostname-, verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

PUT /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

  • Operation ID: putApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: regional-api-proxy (lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM //alb-hostname-, verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

PATCH /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

  • Operation ID: patchApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: regional-api-proxy (lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM //alb-hostname-, verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

DELETE /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

  • Operation ID: deleteApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: regional-api-proxy (lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM //alb-hostname-, verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

HEAD /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

  • Operation ID: headApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: regional-api-proxy (lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM //alb-hostname-, verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

OPTIONS /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

  • Operation ID: optionsApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: regional-api-proxy (lambda/regional-api-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered VPC Lambda. Resolves the internal ALB hostname registered in SSM //alb-hostname-, verifies it is an account-owned internal ALB, signs the HMAC envelope and forwards inside the VPC within 28 s.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

GET /inference/{proxy+}

Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.

  • Operation ID: getInferenceProxy
  • Tags: inference
  • Security: sigv4
  • Integration: aws_proxy, 900 s integration timeout, response streaming

Backend

  • Lambda: inference-streaming-proxy (lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout)
  • Behaviour: Node.js response-streaming VPC Lambda (ROUTING_MODE=regional): the same handler as the global route, reaching the ALB directly instead of through Global Accelerator.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist — inference-proxy
    3. The endpoint's model-server Service in gco-inference (, -canary or -proxy), streamed back through every hop

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
404 Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. —
500 Internal error in the proxy Lambda or the backend. —
502 Bad gateway: the in-cluster proxy or the model server did not answer. —

POST /inference/{proxy+}

Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.

  • Operation ID: postInferenceProxy
  • Tags: inference
  • Security: sigv4
  • Integration: aws_proxy, 900 s integration timeout, response streaming

Backend

  • Lambda: inference-streaming-proxy (lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout)
  • Behaviour: Node.js response-streaming VPC Lambda (ROUTING_MODE=regional): the same handler as the global route, reaching the ALB directly instead of through Global Accelerator.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist — inference-proxy
    3. The endpoint's model-server Service in gco-inference (, -canary or -proxy), streamed back through every hop

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
404 Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. —
500 Internal error in the proxy Lambda or the backend. —
502 Bad gateway: the in-cluster proxy or the model server did not answer. —

HEAD /inference/{proxy+}

Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.

  • Operation ID: headInferenceProxy
  • Tags: inference
  • Security: sigv4
  • Integration: aws_proxy, 900 s integration timeout, response streaming

Backend

  • Lambda: inference-streaming-proxy (lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout)
  • Behaviour: Node.js response-streaming VPC Lambda (ROUTING_MODE=regional): the same handler as the global route, reaching the ALB directly instead of through Global Accelerator.
  • Then:
    1. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    2. inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist — inference-proxy
    3. The endpoint's model-server Service in gco-inference (, -canary or -proxy), streamed back through every hop

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
404 Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. —
500 Internal error in the proxy Lambda or the backend. —
502 Bad gateway: the in-cluster proxy or the model server did not answer. —

Schemas

This api gateway declares no component schemas.