GCO Global API Gateway — API spec sheet¶
API Gateway api-gateway-global · OpenAPI 3.0.3 · API version 1.0.0 · 14 endpoints · 0 schemas
Authenticated global aggregation API for GCO
The single authenticated entry point for GCO. Every request is signed with AWS SigV4 and authorized by IAM; the Lambda proxies behind the routes add the request-bound HMAC envelope that the in-cluster services require, so this API is the only supported caller of the cluster besides the regional bridges. /api/v1/{proxy+} and /inference/{proxy+} reach the regions over Global Accelerator; /api/v1/global/* fans out to every regional API Gateway; /studio/* exists only with the analytics environment.
- Machine-readable document:
docs/openapi/api-gateway-global.json— the documentscripts/generate_api_gateway_openapi.pyreads out of the synthesized CDK stack; this sheet is rendered from it - Interactive console (Swagger UI): https://aws-solutions-library-samples.github.io/global-capacity-orchestrator-on-aws/swagger/api-gateway-global/
- Catalogue index: README.md · interaction diagram
Servers¶
https://{api_id}.execute-api.{region}.{url_suffix}/{stage}
Stage prod of the deployed REST API. The full URL is the ApiEndpoint output of the gco-api-gateway stack.
| Variable | Default | Description |
|---|---|---|
api_id |
"<api-id>" |
REST API id assigned at deploy time. |
region |
"us-east-2" |
deployment_regions.api_gateway in cdk.json. |
stage |
"prod" |
— |
url_suffix |
"amazonaws.com" |
The partition's DNS suffix (${AWS::URLSuffix}). |
Deployment¶
- Endpoint type:
EDGE - Stage:
prod— throttling 1000 req/s (burst 2000), execution loggingINFO, data trace off, metrics on, X-Ray tracing on, access logs on - Conditional routes: operations marked Only when exist under:
analytics: Present only whenanalytics_environment.enabledis true in cdk.json.global-accelerator: Present only in the commercialawspartition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
Resource policy¶
Statements deployed on the REST API (${AWS::AccountId} is the deploying account):
[
{
"Action": "execute-api:Invoke",
"Condition": {
"StringEquals": {
"aws:PrincipalAccount": "${AWS::AccountId}"
}
},
"Effect": "Allow",
"Principal": {
"AWS": "*"
},
"Resource": "execute-api:/*"
},
{
"Action": "execute-api:Invoke",
"Effect": "Allow",
"Principal": {
"AWS": "*"
},
"Resource": "execute-api:/*/GET/studio/*"
}
]
WAF¶
| Priority | Rule | Action | Managed rule group |
|---|---|---|---|
| 0 | PerIPRateLimit |
block | — |
| 1 | NonInferenceBodySizeLimit |
block | — |
| 2 | AWSManagedRulesCommonRuleSet |
group default | AWSManagedRulesCommonRuleSet |
| 3 | AWSManagedRulesKnownBadInputsRuleSet |
group default | AWSManagedRulesKnownBadInputsRuleSet |
| 4 | AWSManagedRulesSQLiRuleSet |
group default | AWSManagedRulesSQLiRuleSet |
| 5 | AWSManagedRulesLinuxRuleSet |
group default | AWSManagedRulesLinuxRuleSet |
| 6 | AWSManagedRulesAmazonIpReputationList |
group default | AWSManagedRulesAmazonIpReputationList |
| 7 | AWSManagedRulesAnonymousIpList |
group default | AWSManagedRulesAnonymousIpList |
Security schemes¶
cognito¶
- Type: apiKey — header
Authorization(cognito_user_pools)
ID token issued by the Studio Cognito user pool, validated by the <project>-studio-cognito-authorizer authorizer.
sigv4¶
- Type: apiKey — header
Authorization(awsSigv4)
AWS Signature Version 4 for service execute-api, signed with IAM credentials of the deploying account that carry execute-api:Invoke on this API (for example awscurl --service execute-api --region <region>).
Endpoints¶
| Method | Path | Summary | Tags |
|---|---|---|---|
GET |
/api/v1/global/health |
Health of every region's cluster | global |
GET |
/api/v1/global/jobs |
List jobs across every region | global |
DELETE |
/api/v1/global/jobs |
Bulk-delete jobs across every region | global |
GET |
/api/v1/global/status |
Aggregated status of every region | global |
GET |
/api/v1/{proxy+} |
Read a control-plane resource | control-plane |
POST |
/api/v1/{proxy+} |
Create a control-plane resource or submit a job | control-plane |
PUT |
/api/v1/{proxy+} |
Replace a control-plane resource | control-plane |
PATCH |
/api/v1/{proxy+} |
Update a control-plane resource | control-plane |
DELETE |
/api/v1/{proxy+} |
Delete a control-plane resource | control-plane |
GET |
/inference/{proxy+} |
Read an inference endpoint's health, info or model list | inference |
POST |
/inference/{proxy+} |
Generate: forward a completion request and stream the response | inference |
HEAD |
/inference/{proxy+} |
Probe an inference endpoint | inference |
GET |
/studio/callback |
OAuth redirect landing page (stub) | studio |
GET |
/studio/login |
Presigned SageMaker Studio login URL | studio |
Endpoint details¶
GET /api/v1/global/health¶
Fans out to /api/v1/health on every regional API Gateway and returns one health object per region.
- Operation ID:
getApiV1GlobalHealth - Tags: global
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
cross-region-aggregator(lambda/cross-region-aggregator/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Discovers every
-regional-api- stack's RegionalApiEndpoint output, fans the request out to each regional API with SigV4 and merges the answers into one response. - Then:
- Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) —
api-gateway-regional - regional-api-proxy VPC Lambda: resolves the ALB from SSM /
/alb-hostname- and signs the HMAC envelope - Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
GET /api/v1/global/jobs¶
Fans out to /api/v1/jobs on every regional API Gateway and merges the results. Requires the regional bridges to be deployed; each region's answer is attributed to it in the response.
- Operation ID:
getApiV1GlobalJobs - Tags: global
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
cross-region-aggregator(lambda/cross-region-aggregator/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Discovers every
-regional-api- stack's RegionalApiEndpoint output, fans the request out to each regional API with SigV4 and merges the answers into one response. - Then:
- Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) —
api-gateway-regional - regional-api-proxy VPC Lambda: resolves the ALB from SSM /
/alb-hostname- and signs the HMAC envelope - Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
DELETE /api/v1/global/jobs¶
Fans out to /api/v1/jobs on every regional API Gateway and merges the results. Requires the regional bridges to be deployed; each region's answer is attributed to it in the response.
- Operation ID:
deleteApiV1GlobalJobs - Tags: global
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
cross-region-aggregator(lambda/cross-region-aggregator/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Discovers every
-regional-api- stack's RegionalApiEndpoint output, fans the request out to each regional API with SigV4 and merges the answers into one response. - Then:
- Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) —
api-gateway-regional - regional-api-proxy VPC Lambda: resolves the ALB from SSM /
/alb-hostname- and signs the HMAC envelope - Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
GET /api/v1/global/status¶
Fans out to /api/v1/status on every regional API Gateway (answered by the manifest-processor) and merges templates, webhooks, resource limits and allowed namespaces per region.
- Operation ID:
getApiV1GlobalStatus - Tags: global
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
cross-region-aggregator(lambda/cross-region-aggregator/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Discovers every
-regional-api- stack's RegionalApiEndpoint output, fans the request out to each regional API with SigV4 and merges the answers into one response. - Then:
- Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) —
api-gateway-regional - regional-api-proxy VPC Lambda: resolves the ALB from SSM /
/alb-hostname- and signs the HMAC envelope - Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
- Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) —
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
GET /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
- Operation ID:
getApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
api-gateway-proxy(lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
- Then:
- AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
POST /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
- Operation ID:
postApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
api-gateway-proxy(lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
- Then:
- AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
PUT /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
- Operation ID:
putApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
api-gateway-proxy(lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
- Then:
- AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
PATCH /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
- Operation ID:
patchApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
api-gateway-proxy(lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
- Then:
- AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
DELETE /api/v1/{proxy+}¶
Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.
Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
- Operation ID:
deleteApiV1Proxy - Tags: control-plane
- Security:
sigv4 - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
api-gateway-proxy(lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout) - Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
- Then:
- AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) —
manifest-processor,health-monitor
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
403 |
Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. |
— |
500 |
Internal error in the proxy Lambda or the backend. | — |
GET /inference/{proxy+}¶
Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.
Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
- Operation ID:
getInferenceProxy - Tags: inference
- Security:
sigv4 - Integration:
aws_proxy, 900 s integration timeout, response streaming
Backend
- Lambda:
inference-streaming-proxy(lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout) - Behaviour: Node.js response-streaming proxy (ROUTING_MODE=global). Accepts GET, HEAD and POST under /inference/
/..., caps the body at 1 MiB, signs the HMAC envelope and streams the model server's response for up to 15 minutes. - Then:
- AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist —
inference-proxy - The endpoint's model-server Service in gco-inference (
, -canary or -proxy), streamed back through every hop
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
404 |
Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
502 |
Bad gateway: the in-cluster proxy or the model server did not answer. | — |
POST /inference/{proxy+}¶
Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.
Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
- Operation ID:
postInferenceProxy - Tags: inference
- Security:
sigv4 - Integration:
aws_proxy, 900 s integration timeout, response streaming
Backend
- Lambda:
inference-streaming-proxy(lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout) - Behaviour: Node.js response-streaming proxy (ROUTING_MODE=global). Accepts GET, HEAD and POST under /inference/
/..., caps the body at 1 MiB, signs the HMAC envelope and streams the model server's response for up to 15 minutes. - Then:
- AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist —
inference-proxy - The endpoint's model-server Service in gco-inference (
, -canary or -proxy), streamed back through every hop
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
404 |
Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
502 |
Bad gateway: the in-cluster proxy or the model server did not answer. | — |
HEAD /inference/{proxy+}¶
Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.
Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.
- Operation ID:
headInferenceProxy - Tags: inference
- Security:
sigv4 - Integration:
aws_proxy, 900 s integration timeout, response streaming
Backend
- Lambda:
inference-streaming-proxy(lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout) - Behaviour: Node.js response-streaming proxy (ROUTING_MODE=global). Accepts GET, HEAD and POST under /inference/
/..., caps the body at 1 MiB, signs the HMAC envelope and streams the model server's response for up to 15 minutes. - Then:
- AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
- Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix —
cluster-gateway - inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist —
inference-proxy - The endpoint's model-server Service in gco-inference (
, -canary or -proxy), streamed back through every hop
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
404 |
Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
502 |
Bad gateway: the in-cluster proxy or the model server did not answer. | — |
GET /studio/callback¶
Unauthenticated MOCK integration returning an empty 200 body: the Cognito hosted-UI OAuth redirect target. The browser redirect carries the authorization code as a query-string parameter and nothing reads the body.
Only when: Present only when analytics_environment.enabled is true in cdk.json.
- Operation ID:
getStudioCallback - Tags: studio
- Security: none (unauthenticated)
- Integration:
mock
Backend
- Integration:
mock - Behaviour: API Gateway answers directly; no Lambda and no backend is involved.
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; empty JSON body. | — |
GET /studio/login¶
Authenticated with a Cognito ID token from the Studio user pool rather than SigV4. Returns a presigned SageMaker Studio URL for the caller's user profile.
Only when: Present only when analytics_environment.enabled is true in cdk.json.
- Operation ID:
getStudioLogin - Tags: studio
- Security:
cognito - Integration:
aws_proxy, 29 s integration timeout
Backend
- Lambda:
analytics-presigned-url(lambda/analytics-presigned-url/handler.py; python3.13, index.handler) - Behaviour: The analytics stack's presigned-URL Lambda (GCOAnalyticsStack): exchanges the Cognito identity for a SageMaker Studio presigned login URL. Synthesized here with an inline stand-in because the analytics stack is not part of this API's template.
Responses
| Status | Description | Content |
|---|---|---|
200 |
Success; the backend's response is returned unchanged. | — |
400 |
Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. |
— |
401 |
Unauthorized: the Cognito ID token is missing, expired or not from the Studio pool. | — |
404 |
Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. | — |
500 |
Internal error in the proxy Lambda or the backend. | — |
Schemas¶
This api gateway declares no component schemas.