diff --git a/README.md b/README.md index c83ba09e8..49eb3f80a 100644 --- a/README.md +++ b/README.md @@ -13,7 +13,7 @@ multiple USSs operating in the same general area to share information while prot privacy. The system is focused on facilitating communication amongst actively operating USSs without details about UAS operations stored in or processed by the DSS. -- [Deploying a DSS instance](https://interuss.github.io/dss) +- [User documentation (including deployment instructions)](https://interuss.github.io/dss) - [Conceptual background on the DSS and services](./concepts.md) - [Introduction to the DSS implementation](./README_DSS.md) - [DSS implementation details](./implementation_details.md) diff --git a/build/README.md b/build/README.md index 565489592..0cf6eb536 100644 --- a/build/README.md +++ b/build/README.md @@ -1,3 +1,3 @@ # Deploying a DSS instance -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/infrastructure/google-manual). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/deployment/). diff --git a/build/deploy/README.md b/build/deploy/README.md index 4f0fd9a56..06af595be 100644 --- a/build/deploy/README.md +++ b/build/deploy/README.md @@ -1,7 +1,7 @@ # Kubernetes deployment via Tanka The documentation and configuration have been moved to the [services directory](../../deploy/services/tanka/). -Architecture, Survivability and Sizing sections have been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/architecture/) +Architecture, Survivability and Sizing sections have been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/background/architecture) ## Migrating configurations to new location diff --git a/deploy/MIGRATION.md b/deploy/MIGRATION.md index 35f0dadd9..fb2926414 100644 --- a/deploy/MIGRATION.md +++ b/deploy/MIGRATION.md @@ -1,3 +1,3 @@ # CockroachDB and Kubernetes version migration -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/operations/migrations). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/operations/crdb-upgrades). diff --git a/deploy/architecture.md b/deploy/architecture.md index a1d2bac7e..ded5dce53 100644 --- a/deploy/architecture.md +++ b/deploy/architecture.md @@ -1,3 +1,3 @@ # Kubernetes deployment -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/architecture). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/background/architecture). diff --git a/deploy/infrastructure/local/minikube/README.md b/deploy/infrastructure/local/minikube/README.md index e9dc30773..b6680a719 100644 --- a/deploy/infrastructure/local/minikube/README.md +++ b/deploy/infrastructure/local/minikube/README.md @@ -1,3 +1,3 @@ # minikube -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/infrastructure/minikube/). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/deployment/infrastructure/minikube/). diff --git a/deploy/infrastructure/modules/terraform-aws-dss/DNS.md b/deploy/infrastructure/modules/terraform-aws-dss/DNS.md index cfd26201a..fea721a95 100644 --- a/deploy/infrastructure/modules/terraform-aws-dss/DNS.md +++ b/deploy/infrastructure/modules/terraform-aws-dss/DNS.md @@ -1,3 +1,3 @@ -# Setup DNS +# terraform-aws-dss -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/infrastructure/aws). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/deployment/infrastructure/aws). diff --git a/deploy/infrastructure/modules/terraform-aws-dss/README.md b/deploy/infrastructure/modules/terraform-aws-dss/README.md index 6607fd185..fea721a95 100644 --- a/deploy/infrastructure/modules/terraform-aws-dss/README.md +++ b/deploy/infrastructure/modules/terraform-aws-dss/README.md @@ -1,3 +1,3 @@ # terraform-aws-dss -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/infrastructure/aws). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/deployment/infrastructure/aws). diff --git a/deploy/infrastructure/modules/terraform-google-dss/DNS.md b/deploy/infrastructure/modules/terraform-google-dss/DNS.md index 2e25fb0b3..3b7953f00 100644 --- a/deploy/infrastructure/modules/terraform-google-dss/DNS.md +++ b/deploy/infrastructure/modules/terraform-google-dss/DNS.md @@ -1,3 +1,3 @@ -# Setup DNS +# terraform-google-dss -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/infrastructure/google). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/deployment/infrastructure/google). diff --git a/deploy/infrastructure/modules/terraform-google-dss/README.md b/deploy/infrastructure/modules/terraform-google-dss/README.md index 108289c86..3b7953f00 100644 --- a/deploy/infrastructure/modules/terraform-google-dss/README.md +++ b/deploy/infrastructure/modules/terraform-google-dss/README.md @@ -1,3 +1,3 @@ # terraform-google-dss -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/infrastructure/google). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/deployment/infrastructure/google). diff --git a/deploy/operations/certificates-management/README.md b/deploy/operations/certificates-management/README.md index 20c9b1253..289c8d4ea 100644 --- a/deploy/operations/certificates-management/README.md +++ b/deploy/operations/certificates-management/README.md @@ -1,3 +1,3 @@ # Certificates management -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/operations/certificates-management). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/deployment/pooling/yugabyte). diff --git a/deploy/operations/pooling-crdb.md b/deploy/operations/pooling-crdb.md index 25b50664d..7338625bf 100644 --- a/deploy/operations/pooling-crdb.md +++ b/deploy/operations/pooling-crdb.md @@ -1,3 +1,3 @@ # DSS Pooling -This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/operations/pooling-crdb). +This documentation has been moved to [interuss.github.io/dss](https://interuss.github.io/dss/dev/background/pooling-crdb). diff --git a/deploy/services/helm-charts/dss/README.md b/deploy/services/helm-charts/dss/README.md index 4b8405c63..89d18a38a 100644 --- a/deploy/services/helm-charts/dss/README.md +++ b/deploy/services/helm-charts/dss/README.md @@ -3,11 +3,11 @@ This directory provides an [Helm Chart](https://helm.sh/) to deploy the DSS and ## Requirements 1. A Kubernetes cluster should be running and you should be properly authenticated. Requirements and instructions to create a new Kubernetes cluster can be found here: - * [AWS](../../../../docs/infrastructure/aws.md) - * [Google](../../../../docs/infrastructure/google.md) - * [Minikube](../../../../docs/infrastructure/minikube.md) + * [AWS](../../../../docs/deployment/infrastructure/aws.md) + * [Google](../../../../docs/deployment/infrastructure/google.md) + * [Minikube](../../../../docs/deployment/infrastructure/minikube.md) -2. Create the certificates and apply them to the cluster using the instructions [here](../../../../docs/operations/certificates-management.md) +2. Create the certificates and apply them to the cluster using the instructions [here](../../../../docs/deployment/pooling/index.md) 3. Install [Helm](https://helm.sh/) version 3.11.3 or higher ## Usage diff --git a/deploy/services/tanka/README.md b/deploy/services/tanka/README.md index bc477e287..2c03a7365 100644 --- a/deploy/services/tanka/README.md +++ b/deploy/services/tanka/README.md @@ -5,11 +5,11 @@ This folder contains a set of configuration to be used with [tanka](https://tank ## Requirements 1. A Kubernetes cluster should be running and you should be properly authenticated. Requirements and instructions to create a new Kubernetes cluster can be found here: - * [AWS](../../../docs/infrastructure/aws.md) - * [Google](../../../docs/infrastructure/google.md) - * [Minikube](../../../docs/infrastructure/minikube.md) + * [AWS](../../../docs/deployment/infrastructure/aws.md) + * [Google](../../../docs/deployment/infrastructure/google.md) + * [Minikube](../../../docs/deployment/infrastructure/minikube.md) -2. Create the certificates and apply them to the cluster using the instructions [here](../../../docs/operations/certificates-management.md) +2. Create the certificates and apply them to the cluster using the instructions [here](../../../docs/deployment/pooling/index.md) 3. Install [Tanka](https://tanka.dev/install) ## Usage diff --git a/docs/.nav.yml b/docs/.nav.yml index ccbd43642..9caaee967 100644 --- a/docs/.nav.yml +++ b/docs/.nav.yml @@ -1,6 +1,6 @@ nav: - "Getting Started": index.md - - "Architecture": architecture/ - - "Deploy a DSS instance to": infrastructure/ + - "Deploy a DSS instance": deployment/ - "Operate a DSS instance": operations/ - - "Deployment checklist": deployment_checklist.md + - "Decommission a DSS instance": decommissioning/ + - "Background": background/ diff --git a/docs/TODO.txt b/docs/TODO.txt new file mode 100644 index 000000000..2bd17d7f6 --- /dev/null +++ b/docs/TODO.txt @@ -0,0 +1,9 @@ +Remaining documentation enhancement tasks: + +* Ensure deployment/pooling instructions are accurate, especially regarding the expected working directory when calling make_certs.py +* When linking to the repository, ensure repo version matches documentation version (e.g., 0.22.0 tag) +* Ensure VAR_* references outside the manual deployment instructions are defined for the terraform approach or harmonized +* In Minikube documentation, replace `dss-local-cluster` with `${cluster_name}` or similar and have users export/set that once initially +* Ensure Minikube documentation describes how to accomplish pooling +* Actually identify which tools are needed in operations/index.md rather than "some of these tools" -- the purpose of this statement is procedural to ensure ops personnel have the right tools installed, so we should identify which tools those are rather than having them guess or figure it out themselves +* Add/migrate helm documentation if that will be an officially-supported way to deploy services diff --git a/docs/architecture/.nav.yml b/docs/architecture/.nav.yml deleted file mode 100644 index baa69923d..000000000 --- a/docs/architecture/.nav.yml +++ /dev/null @@ -1,4 +0,0 @@ -flatten_single_child_sections: true -nav: - - "Introduction": index.md - - "Sizing": sizing.md diff --git a/docs/assets/create_pool_2.puml b/docs/assets/create_pool_2.puml deleted file mode 100644 index 34c3cc917..000000000 --- a/docs/assets/create_pool_2.puml +++ /dev/null @@ -1,33 +0,0 @@ -'To render with PlantUML: -' java -jar plantuml.jar -o generated create_pool_2.puml -@startuml -skinparam dpi 300 -participant "USS 1" as USS1 -participant "USS 2" as USS2 -participant "Pool state" as PoolState - -note over USS1: Create DSS instance,\ninitialize cluster -note over PoolState: Pool ready with\n1 instance -note over USS1: Run prober on DSS instance\nto verify functionality -note over PoolState: Pool verified with\n1 instance - -USS2 -> USS1: Request ca.crt and\nCRDB node addresses -note over USS2: Follow instructions\nto make-certs.py -USS1 --> USS2: Provide ca.crt and\nCRDB node addresses -note over USS2: Run make-certs.py -USS2 -> USS1: Provide combined ca.crt -note over USS2: Follow instructions\nto `tk apply` -note over USS1: Restart CRDB nodes\nwith combined ca.crt -note over PoolState: Pool ready to accept\nsecond instance -USS1 -> USS2: Verify ready for pool -note over USS2: `tk apply` to deploy\nDSS instance -note over PoolState: Pool ready with\n2 instances -note over USS2: Run prober on DSS instance\nto verify functionality -note over USS1: Run prober on DSS instance\nto verify no regression -note right of USS1: USS 1 and/or USS 2 run\ninterop test on DSS instances\nto verify functionality -note over PoolState: Pool verified with\n2 instances - -USS2 -> USS1: Provide CRDB node addresses -note over USS1: Update JoinExisting\nnode list -note over PoolState: USS 1 will automatically\nrejoin pool upon restart -@enduml diff --git a/docs/assets/create_pool_n.gv b/docs/assets/create_pool_n.gv deleted file mode 100644 index 069450e0a..000000000 --- a/docs/assets/create_pool_n.gv +++ /dev/null @@ -1,55 +0,0 @@ -// To render: -// dot -Tpng -ogenerated/create_pool_n.png create_pool_n.gv -digraph { - node [shape=box] - graph [dpi = 300]; - - Legendn [label="Instance n",color=red] - PS1 [label="Pool verified with\nn-1 instances",shape=note] - Legend1 [label="Instance 1",color=green] - Legend2 [label="Instance i",color=blue] - Legendnm1 [label="Instance n-1",color=yellow] - USSn_1a [label="Request ca.crt and\nCRDB node addresses\nfrom all USSs",color=red] - USSn_1b [label="Follow instructions\nto make-certs.py",color=red] - PS1 -> USSn_1a - PS1 -> USSn_1b -> USSn_2 - USS1_1 [label="Provide ca.crt and\nCRDB node addresses",color=green] - USS2_1 [label="Provide ca.crt and\nCRDB node addresses",color=blue] - USSnm1_1 [label="Provide ca.crt and\nCRDB node addresses",color=yellow] - USSn_1a -> USS1_1 - USSn_1a -> USS2_1 - USSn_1a -> USSnm1_1 - USSn_2 [label="Run make-certs.py",color=red] - USSn_3 [label="Provide combined ca.crt\nto all USSs",color=red] - USS2_1 -> USSn_2 -> USSn_3 - USS1_2 [label="Restart CRDB\nnodes with\ncombined ca.crt",color=green] - USS2_2 [label="Restart CRDB\nnodes with\ncombined ca.crt",color=blue] - USSnm1_2 [label="Restart CRDB\nnodes with\ncombined ca.crt",color=yellow] - USSn_4 [label="Follow instructions\nto `tk apply`",color=red] - USSn_2 -> USSn_4 - USS1_1 -> USSn_4 - USS2_1 -> USSn_4 - USSnm1_1 -> USSn_4 -> USSn_5 - USSn_3 -> USS1_2 -> USSn_5 - USSn_3 -> USS2_2 -> USSn_5 - USSn_3 -> USSnm1_2 -> USSn_5 - USSn_5 [label="`tk apply` to deploy\nDSS instance",color=red] - USSn_6 [label="Run prober on DSS\ninstance to\nverify functionality",color=red] - USS1_4 [label="Run prober on DSS\ninstance to\nverify no regression",color=green] - USS2_4 [label="Run prober on DSS\ninstance to\nverify no regression",color=blue] - USSnm1_4 [label="Run prober on DSS\ninstance to\nverify no regression",color=yellow] - USSn_5 -> USSn_6 -> USS_1 -> PS4 - USSn_5 -> USS1_4 -> USS_1 -> USSn_7 - USSn_5 -> USS2_4 -> USS_1 - USSn_5 -> USSnm1_4 -> USS_1 - USS_1 [label="Any USS run\ninterop test on DSS instances\nto verify functionality",color=magenta] - PS4 [label="Pool verified with\nn instances",shape=note] - USSn_7 [label="Provide CRDB node\naddresses to all USSs",color=red] - USS1_6 [label="Update JoinExisting\nnode list",color=green] - USS2_6 [label="Update JoinExisting\nnode list",color=blue] - USSnm1_6 [label="Update JoinExisting\nnode list",color=yellow] - USSn_7 -> USS1_6 -> PS5 - USSn_7 -> USS2_6 -> PS5 - USSn_7 -> USSnm1_6 -> PS5 - PS5 [label="All instances will automatically\nrejoin pool upon restart",shape=note] -} diff --git a/docs/assets/deployment_layers.gv b/docs/assets/deployment_layers.gv deleted file mode 100644 index 65afae534..000000000 --- a/docs/assets/deployment_layers.gv +++ /dev/null @@ -1,11 +0,0 @@ -digraph { - node [shape=box,fillcolor=white,style=filled] - graph [dpi = 300]; - - Start -> Infrastructure -> Services -> Operations - - Start [shape=circle] - Infrastructure [label="Infrastructure\n(VMs, Kubernetes cluster, IP addresses, DNS resolution, etc)"] - Services [label="Services\n(Deployment of jobs to Kubernetes cluster)"] - Operations [label="Operations\n(Ongoing maintenance activities like pooling, upgrading, etc)"] -} diff --git a/docs/assets/generated/create_pool_2.png b/docs/assets/generated/create_pool_2.png deleted file mode 100644 index 08dba0655..000000000 Binary files a/docs/assets/generated/create_pool_2.png and /dev/null differ diff --git a/docs/assets/generated/create_pool_n.png b/docs/assets/generated/create_pool_n.png deleted file mode 100644 index d13212bc1..000000000 Binary files a/docs/assets/generated/create_pool_n.png and /dev/null differ diff --git a/docs/assets/generated/deployment_layers.png b/docs/assets/generated/deployment_layers.png deleted file mode 100644 index f040ad2ce..000000000 Binary files a/docs/assets/generated/deployment_layers.png and /dev/null differ diff --git a/docs/background/.nav.yml b/docs/background/.nav.yml new file mode 100644 index 000000000..2bdde7773 --- /dev/null +++ b/docs/background/.nav.yml @@ -0,0 +1,8 @@ +nav: + - "Overview": index.md + - "Architecture": architecture.md + - "Authentication": authentication.md + - "Sizing": sizing.md + - "Pooling": pooling.md + - "Pooling: YugabyteDB specifics": pooling-yugabyte.md + - "Pooling: CRDB specifics": pooling-crdb.md diff --git a/docs/architecture/index.md b/docs/background/architecture.md similarity index 94% rename from docs/architecture/index.md rename to docs/background/architecture.md index a71567d11..476a5cccc 100644 --- a/docs/architecture/index.md +++ b/docs/background/architecture.md @@ -1,4 +1,4 @@ -# Architecture +# Deployment architecture ## Introduction @@ -30,20 +30,6 @@ Top level simplified view, with one replica shown and yugabyte services regroupe To reduce the number of required public load balancers, we do use an intermediate reverse proxy to expose the ports of Yugabyte master and tserver on a shared public IP per stateful set instance. Usual Kubernetes load balancers can't assign connection based on ports out of the box, so we use the reverse proxy to dispatch connections on both services depending on the connected port. -### Terminology notes - -See [teminology notes](../operations/pooling.md#terminology-notes). - -## Pooling - -### Objective - -See [Pooling Objective](../operations/pooling.md#objective) and subsections. - -### Additional requirements - -See [Additional requirements](../operations/pooling.md#additional-requirements). - ### Survivability One of the primary design considerations of the DSS is to be very resilient to diff --git a/docs/operations/authentication.md b/docs/background/authentication.md similarity index 100% rename from docs/operations/authentication.md rename to docs/background/authentication.md diff --git a/docs/background/index.md b/docs/background/index.md new file mode 100644 index 000000000..c16b59e0e --- /dev/null +++ b/docs/background/index.md @@ -0,0 +1,10 @@ +# Background overview + +Background information on various aspects of the InterUSS DSS product may be found here. + +* [Architecture](./architecture.md) +* [Authentication](./authentication.md) +* [Sizing](./sizing.md) +* [Pooling](./pooling.md) + * [With CockroachDB](./pooling-crdb.md) + * [With YugabyteDB](./pooling-yugabyte.md) diff --git a/docs/operations/pooling-crdb.md b/docs/background/pooling-crdb.md similarity index 51% rename from docs/operations/pooling-crdb.md rename to docs/background/pooling-crdb.md index 4e24e5fab..ae3e81ce9 100644 --- a/docs/operations/pooling-crdb.md +++ b/docs/background/pooling-crdb.md @@ -1,50 +1,9 @@ -# DSS Pooling (CockroachDB) +# DSS Pooling using CockroachDB -!!! note - This document is about pooling with **CockroachDB**. Yugabyte documentation is [there](./pooling.md). +This document describes the specifics of using CockroachDB as the data store to +accomplish [pooling](./pooling.md). -## Introduction - -The DSS is designed to be deployed in a federated manner where multiple -organizations each host a DSS instance, and all of those instances interoperate. -Specifically, if a change is made on one DSS instance, that change may be read -from a different DSS instance. A set of interoperable DSS instances is called a -"pool", and the purpose of this document is to describe how to form and maintain -a DSS pool. - -It is expected that there will be exactly one production DSS pool for any given -DSS region, and that a DSS region will generally match aviation jurisdictional -boundaries (usually national boundaries). A given DSS region (e.g., -Switzerland) will likely have one pool for production operations, and an -additional pool for partner qualification and testing (per, e.g., -F3411-19 A2.6.2). - -### Terminology notes - -CockroachDB (CRDB) establishes a distributed data store called a "cluster". -This cluster stores the DSS Airspace Representation (DAR) in multiple SQL -databases within that cluster. This cluster is composed of many CRDB nodes, -potentially hosted by multiple organizations. - -Kubernetes manages a set of services in a "cluster". This is an entirely -different thing from the CRDB data store, and this type of cluster is what the -deployment instructions refer to. A Kubernetes cluster contains one or more -node pools: collections of machines available to run jobs. This node pool is an -entirely different thing from a DSS pool. - -## Objective - -A pool of InterUSS-compatible DSS instances is established when all of the -following requirements are met: - -1. Each CockroachDB node is addressable by every other CockroachDB node -1. Each CockroachDB node is discoverable -1. Each CockroachDB node accepts the certificates of every other node -1. The CockroachDB cluster is initialized - -The procedures in this document are intended to achieve all the objectives -listed above, but these procedures are not the only ways to achieve the -objectives. +## How CockroachDB meets pooling requirements ### "Each CockroachDB node is addressable by every other CockroachDB node" @@ -132,32 +91,6 @@ following those instructions. [Google's Public NTP](https://developers.google.com/time/) when running in a multi-cloud environment. -## Creating a new pool -Although all DSS instances are equal peers, one DSS instance must be chosen to -create the pool initially. After the pool is established, one additional DSS -instance can join it. After that joining process is complete, it can be -repeated any number of times to add additional DSS instances, though 7 is the -maximum recommended number of DSS instances for performance reasons. The -following diagram illustrates the pooling process for the first two instances: - -![DSS pooling for first two participants](../assets/generated/create_pool_2.png) - -The nth instance joins in almost the same way as the second instance; the -diagram below illustrates action dependencies between the existing and joining -DSS instances to allow the new instance to join the existing pool: - -![DSS pooling for nth participants](../assets/generated/create_pool_n.png) - -### Establishing a pool with first instance -The USS owning the first DSS instance should follow -[the deployment instructions](index.md). They are not joining any existing -cluster, and specifically `VAR_SHOULD_INIT` should be set `true` to initialize -the CRDB cluster. Upon deployment completion, the following should be run against the DSS instance to verify functionality: - - - The [prober test](https://github.com/interuss/monitoring/blob/main/monitoring/prober/README.md) - - The [USS qualifier](https://github.com/interuss/monitoring/tree/main/monitoring/uss_qualifier), using the [DSS Probing](https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/configurations/dev/dss_probing.yaml) configuration - - ### Joining an existing pool with new instance A USS wishing to join an existing pool (of perhaps just one instance following the prior section) should follow [the deployment instructions](index.md). They @@ -183,56 +116,3 @@ on the full pool (including the newly-added instance). Finally, the joining USS should provide its node addresses to all other participants in the pool, and each other participant should add those addresses to the list of CRDB nodes their CRDB nodes will attempt to contact upon restart. - -## Leaving a pool - -In an event that requires removing CockroachDB nodes we need to properly and -safely decommission to reduce risks of outages. - -It is never a good idea to take down more than half the number of nodes -available in your cluster as doing so would break quorum. If you need to take -down that many nodes please do it in smaller steps. - -Note: If you are removing a specific node in a Statefulset, please know that -Kubernetes does not support removal of specific node; it automatically -re-creates the node if you delete it with `kubectl delete pod`. You will need -to scale down the Statefulset and that removes the last node first (ex: -`cockroachdb-n` where `n` is the `size of statefulset - 1`, `n` starts at 0) - -1. Check if all nodes are healthy and there are no - under-replicated/unavailable ranges: - - `kubectl exec -it cockroachdb-0 -- cockroach node status --ranges --certs-dir=cockroach-certs/` - - 1. If there are under-replicated ranges changes are it is because of a node - failure. If all nodes are healthy then it should auto recover. - - 1. If there are unhealthy nodes please investigate and fix them so that the - ranges can return to a healthy state - -1. Identify the node id we intend to decommission from the previous commands - then decommission them. The following command assumes that `cockroachdb-0` is - not targeted for decommission otherwise select a different instance to - connect to: - - `kubectl exec -it cockroachdb-0 -- cockroach node decommission [ ...] --certs-dir=cockroach-certs/` - -1. If the command executes successfully all targeted nodes should not host any - ranges. Repeat step one to verify - - a. In the event of a hung decommission please recommission the nodes and - repeat the above step with smaller number of nodes to decommission: - - `kubectl exec -it cockroachdb-0 -- cockroach node recommission [ ...] --certs-dir=cockroach-certs/` - -1. Power down the pods or delete the Statefulset, whichever is applicable - - a. Again, Statefulsets does not support deleting specific pods, as it will - restart it immediately you will need to scale down understanding that it - will remove node `cockroachdb-n` first; where `n` is the - `size of statefulset - 1`. - - To proceed: `kubectl scale statefulset cockroachdb --replicas=` - - b. To remove the entire Statefulset: - `kubectl delete statefulset cockroachdb` diff --git a/docs/background/pooling-yugabyte.md b/docs/background/pooling-yugabyte.md new file mode 100644 index 000000000..fc0d67ed5 --- /dev/null +++ b/docs/background/pooling-yugabyte.md @@ -0,0 +1,129 @@ +# DSS Pooling using YugabyteDB + +This document describes the specifics of using YugabyteDB as the data store to +accomplish [pooling](./pooling.md). + +## How Yugabyte meets pooling requirements + +### "Each Yugabyte node is addressable by every other Yugabyte node" + +Every Yugabyte node must have its own externally-accessible hostname (e.g., +1.tserver.db.dss-prod.example.com), or its own hostname:port combination (e.g., +db.dss-prod.example.com:26258). + +There are two type of nodes in a Yugabyte cluster: Master and TServer. Both ones +must be accessible. The ports on which Yugabyte communicates must be open to +others participants: + +* Master: gRPC: **7100** +* TServer: gRPC: **9100** +* Master: Admin UI: 7000 +* TServer: Admin UI: 9000 +* TServer: ycql: 9042 +* TServer: ysql: 5433 +* TServer: metrics: 13000 +* TServer: metrics: 12000 + +The ports in bold are mandatory. The others ones are needed for management UI, +the UI won't work correctly if any of those port is not reachable by other +nodes, except on the master node. + +!!! info + The Helm charts and the Tanka files only expose mandatory ports as they are the + only ones secure. If usage of the UI is needed in a pool with multiple + participants, you must find a way to open those ports in a way secure enough + for your deployments. + Most of those non-mandatory ports do not offer authentication nor encryption + (or confidentiality). A secure method is required, such as an Istio mesh or + a local private network. + +This requirement may be verified by conducting a standard TLS diagnostic +(like [this one](https://www.wormly.com/test_ssl)) on the hostname:port +for each TServer node (e.g., 0.tserver.db.dss.example.com:7100). The "Trust" +characteristic will not pass because the certificate is issued by +a custom CA which is not a generally-trusted root CA, but we +explicitly enable trust by manually exchanging the trusted CA public keys +in ca.crt (see "Each Yugabyte node accepts the certificates of every other +node" below). However, all other checks should generally pass. + +!!! danger + It's recommended to restrict access to all ports and only allow IPs of + others participants. However, guides and deployment tooling haven't been adapted yet. + +### "Each Yugabyte node is discoverable" + +When a Yugabyte node is brought online, it must know how to connect to the +existing network of nodes. This is accomplished by providing an explicit list of +nodes to contact. Each node contacted will provide a list of nodes it is +connected to in the network ("gossip"), so not every node must be present in the +explicit list, but it's recommended to do so. The explicit list shall contain, +at a minimum, all known nodes when creating the DSS instance and shall be +updated regularly. Yugabyte nodes have some difficulites to locate primary nodes +if they don't have the full list of known nodes. + +### "Each Yugabyte node accepts the certificates of every other node" + +Yugabyte uses TLS to secure connections, and TLS includes a mechanism to +ensure the identity of the server being contacted. This mechanism requires a +trusted root Certificate Authority (CA) to sign a certificate containing the +public key of a particular server, so a client connecting to that server can +verify that the certificate (containing the public key) is endorsed by the root +CA as being genuine. Yugabyte certificates require a claim that standard web CAs +will not sign, so instead each USS acts as their own root CA. When USS 1 +is presented with certificates signed by USS 2's CA, USS 1 must know that it +can trust that certificate. We accomplish this by exchanging all USSs' CA +public keys out-of-band in ca.crt, and specifying that certificates signed by +any of the public keys in ca.crt should be accepted when considering the +validity of certificates presented to establish a TLS connection between nodes. + +The private CA key `dss-certs.py` generates is stored in the `ca` folder. The +private CA key is used to generate all node certificates and client +certificates. Once a pool is established, a USS avoids regenerating this CA +keypair, and use the existing ones by default. If a USS generates a new CA +keypair, the new public key must be added to the pool's combined ca.crt, and all +USSs in the pool must adopt the new combined ca.crt before any nodes using +certificates generated by the new CA private key will be accepted by the pool. + +### "The Yugabyte cluster is initialized" + +A Yugabyte cluster of databases is like the Ship of Theseus: it is composed of +many nodes which may all be replaced, one by one, so that a given Yugabyte +cluster eventually contains none of its original nodes. Unlike the Ship of +Theseus, however, a cluster is clearly identified by its cluster ID (e.g., +b2537de3-166f-42c4-aae1-742e094b8349) -- if the cluster ID is the same, it is +the same cluster (and vice versa). Once for the entire lifetime of the Yugabyte +cluster, it is created automatically during initialization of the first set of +Yugabyte nodes, if all initial nodes see each others in a uninitialized state. + +## Additional requirements + +These requirements must be met by every DSS instance joining an +InterUSS-compatible pool. The deployment instructions produce a system that +complies with all these requirements, so this section may be ignored if +following those instructions. + +- All Yugabyte nodes must be run in secure mode. + - use_node_to_node_encryption enabled + - use_client_to_server_encryption enabled + - node_to_node_encryption_use_client_certificates enabled + - allow_insecure_connections disabled +- All DSS instances in the same cluster must point their ntpd at the same NTP + Servers. + +## Placement + +It's important to maintain a good placement strategy, ensuring data availability +in case of failures. + +We do recommend a minimum of 3 participants and one copy in each participants. + +You may use the `modify_placement_info` command to set placement settings. +Example: + +* ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 modify_placement_info dss.uss-1,dss.uss-2,dss.uss-3 3`` +* ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 modify_placement_info dss.uss-1,dss.uss-2,dss.uss-3,dss.uss-4,dss.uss-5 5`` + +You may however use a different strategy depending on your availability needs, +e.g. you may want to avoid common datacenter between the same DSS instance. To +do so, define a strategy in your pool and edit placement information as needed. +More information is available in [Yugabyte documentation](https://docs.yugabyte.com/preview/admin/yb-admin/#modify-placement-info). diff --git a/docs/background/pooling.md b/docs/background/pooling.md new file mode 100644 index 000000000..b4b38a9df --- /dev/null +++ b/docs/background/pooling.md @@ -0,0 +1,75 @@ +# DSS Pooling + +This page provides background information on what pooling is and how it is +accomplished. For pooling operational instructions, see +[deployment](../deployment/index.md). + +## Introduction + +The DSS is designed to be deployed in a federated manner where multiple +organizations each host a DSS instance, and all of those instances interoperate. +Specifically, if a change is made on one DSS instance, that change may be read +from a different DSS instance. A set of interoperable DSS instances is called a +"pool", and the purpose of this document is to describe how to form and maintain +a DSS pool. + +It is expected that there will be exactly one production DSS pool for any given +DSS region, and that a DSS region will generally match aviation jurisdictional +boundaries (usually national boundaries). A given DSS region (e.g., +Switzerland) will likely have one pool for production operations, and an +additional pool for partner qualification and testing (per, e.g., +F3411-19 A2.6.2). + +### Terminology notes + +Some databases establish a distributed data store called a "cluster". This cluster +stores the DSS Airspace Representation (DAR) in multiple SQL databases within +that cluster. This cluster is generally of many nodes of the database, potentially +hosted by multiple organizations. + +Kubernetes manages a set of services in a "cluster". This is an entirely +different thing from the database cluster, and this type of cluster is what +the deployment instructions refer to. A Kubernetes cluster contains one or more +node pools: collections of machines available to run jobs. This node pool is an +entirely different thing from a DSS pool. + +## Objective + +A pool of InterUSS-compatible DSS instances is established when all of the +following requirements are met: + +1. Each database node is addressable by every other database node +1. Each database node is discoverable +1. Each database node accepts the certificates of every other node +1. The database cluster is initialized + +These requirements are met differently depending on the data store used: +- [YugabyteDB](./pooling-yugabyte.md) +- [CockroachDB](./pooling-crdb.md) + +## Creating a new pool +All DSS instances are equal peers, and any set of DSS instance can be chosen to +create the pool initially. After the pool is established, additional DSS +instance can join it. After that joining process is complete, it can be +repeated any number of times to add additional DSS instances, though 7 is the +maximum recommended number of DSS instances for performance reasons. The +following diagram illustrates the pooling process for the first two instances: + +![DSS pooling as a first, alone participant](../assets/generated/pool_new_1.png) + + +![DSS pooling with first 3 first participants](../assets/generated/pool_new_3.png) + +Adding participant is illustrated below. Some actions marked with `(once)` need +to be run only once by one participant otherwise all participants in the current +pool must ran then. + +![DSS pooling with new participant](../assets/generated/pool_add.png) + +### Establishing a pool with the first instance +The USSs owning the first DSS instances should follow +[the deployment instructions](../deployment/index.md). + +### Joining an existing pool with new instance +A USS wishing to join an existing pool (of perhaps just one instance following +the prior section) should follow [the deployment instructions](index.md). diff --git a/docs/architecture/sizing.md b/docs/background/sizing.md similarity index 100% rename from docs/architecture/sizing.md rename to docs/background/sizing.md diff --git a/docs/decommissioning/.nav.yml b/docs/decommissioning/.nav.yml new file mode 100644 index 000000000..a45ffe19a --- /dev/null +++ b/docs/decommissioning/.nav.yml @@ -0,0 +1,6 @@ +nav: + - "Overview": index.md + - "Leaving a CockroachDB pool": depooling-crdb.md + - "Leaving a YugabyteDB pool": depooling-yugabyte.md + - "Terraform": terraform.md + - "Minikube": minikube.md diff --git a/docs/decommissioning/depooling-crdb.md b/docs/decommissioning/depooling-crdb.md new file mode 100644 index 000000000..349d6e80e --- /dev/null +++ b/docs/decommissioning/depooling-crdb.md @@ -0,0 +1,46 @@ +## Leaving a pool + +In an event that requires removing CockroachDB nodes we need to properly and +safely decommission to reduce risks of outages. + +It is never a good idea to take down more than half the number of nodes +available in your cluster as doing so would break quorum. If you need to take +down that many nodes please do it in smaller steps. + +Note: If you are removing a specific node in a Statefulset, please know that +Kubernetes does not support removal of specific node; it automatically +re-creates the node if you delete it with `kubectl delete pod`. You will need +to scale down the Statefulset and that removes the last node first (ex: +`cockroachdb-n` where `n` is the `size of statefulset - 1`, `n` starts at 0) + +1. Check if all nodes are healthy and there are no + under-replicated/unavailable ranges:
`kubectl exec -it cockroachdb-0 -- cockroach node status --ranges --certs-dir=cockroach-certs/` + + 1. If there are under-replicated ranges changes are it is because of a node + failure. If all nodes are healthy then it should auto recover. + + 1. If there are unhealthy nodes please investigate and fix them so that the + ranges can return to a healthy state + +1. Identify the node id we intend to decommission from the previous commands + then decommission them. The following command assumes that `cockroachdb-0` is + not targeted for decommission otherwise select a different instance to + connect to:
`kubectl exec -it cockroachdb-0 -- cockroach node decommission [ ...] --certs-dir=cockroach-certs/` + +1. If the command executes successfully all targeted nodes should not host any + ranges. Repeat step one to verify + + a. In the event of a hung decommission please recommission the nodes and + repeat the above step with smaller number of nodes to decommission:
`kubectl exec -it cockroachdb-0 -- cockroach node recommission [ ...] --certs-dir=cockroach-certs/` + +1. Power down the pods or delete the Statefulset, whichever is applicable + + a. Again, Statefulsets does not support deleting specific pods, as it will + restart it immediately you will need to scale down understanding that it + will remove node `cockroachdb-n` first; where `n` is the + `size of statefulset - 1`. + + To proceed: `kubectl scale statefulset cockroachdb --replicas=` + + b. To remove the entire Statefulset: + `kubectl delete statefulset cockroachdb` diff --git a/docs/decommissioning/depooling-yugabyte.md b/docs/decommissioning/depooling-yugabyte.md new file mode 100644 index 000000000..b49a1a3b3 --- /dev/null +++ b/docs/decommissioning/depooling-yugabyte.md @@ -0,0 +1,79 @@ +## Leaving a pool + +In an event that requires removing Yugabyte nodes we need to properly and +safely decommission to reduce risks of outages. + +It is never a good idea to take down more than half the number of nodes +available in your cluster as doing so would break quorum. If you need to take +down that many nodes please do it in smaller steps. + +Ensure placement info is how you want it after removal. Ensure you're not +requesting impossible placement by removing nodes, otherwise it won't be +possible to request node deletion. See the section below for placement +requirements. + +Note: If you are removing a specific node in a Statefulset, please know that +Kubernetes does not support removal of specific node; it automatically +re-creates the node if you delete it with `kubectl delete pod`. You will need +to scale down the Statefulset and that removes the last node first (ex: +`yb-tserver-n` where `n` is the `size of statefulset - 1`, `n` starts at 0) + +1. Check if all nodes are healthy in the web ui. + +1. Connect to a Yugabyte master and copy certs, like introduced in the previous + section. + +1. For each TServer node to be removed: + + 1. Blacklist one node in your cluster. + + ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 change_blacklist ADD [TSERVER_PUBLIC_HOSTNAME]`` + + 1. Wait for the node to be drained (no user tablet-peer or system-table-peer + in the gui). If node is not draining, you may have placement constraints + that prevent the removal of the node. + + 1. Stop one node in your cluster. + + 1. Wait until the node is marked as down and cluster will go into a + non-healthy state then wait for recovery. When everything is green again + proceed. Depending on settings, it may take time (15m) before the node is + marked as dead. + + 1. Remove the node: + + ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 remove_tablet_server [TSERVER_ID]`` + + If the command is giving you an error, data of the node may not have been + drain correctly dues to placement constraints. + + 1. Remove the node from the black list: + + ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 change_blacklist REMOVE [TSERVER_PUBLIC_HOSTNAME]`` + + 1. Fully remove the node in your cluster. + + E.g you may delete persistent volumes. + + +1. For each Master node to be removed: + + 1. Remove the master from the master list + + ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 change_master_config REMOVE_SERVER [PUBLIC HOSTNAME] 7100`` + + If the master node to be removed is the current leader, you may make it step + down with the following command: + + ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 master_leader_stepdown`` + +Finally, each pool participant should remove master addresses from the +`yugabyte_external_nodes` list their Yugabyte nodes will attempt to contact upon +restart and remove the CA of the participant. + +!!! note + Quick reminder for CA management: + + Remove the old CA, use `./dss-certs.sh remove-pool-ca ` + Finally, apply certificates on the kubernetes cluster with + `./dss-certs.sh apply` diff --git a/docs/decommissioning/index.md b/docs/decommissioning/index.md new file mode 100644 index 000000000..7c8526514 --- /dev/null +++ b/docs/decommissioning/index.md @@ -0,0 +1,11 @@ +# Decommissioning a DSS instance + +Before infrastructure decommissioning, the DSS instance must gracefully leave the pool or their pool components will be considered down permanently, leading to difficulty achieving quorum. + +* [Leaving a CockroachDB pool](depooling-crdb.md) +* [Leaving a YugabyteDB pool](depooling-yugabyte.md) + +Final decommissioning depends on how the DSS was deployed: + +* [Terraform](./terraform.md) +* [Minikube](./minikube.md) diff --git a/docs/decommissioning/minikube.md b/docs/decommissioning/minikube.md new file mode 100644 index 000000000..0bb176751 --- /dev/null +++ b/docs/decommissioning/minikube.md @@ -0,0 +1,5 @@ +# Decommissioning a Minikube DSS deployment + +## Clean up + +To delete all resources, run `minikube delete -p dss-local-cluster`. Note that this operation can't be reverted and all data will be lost. diff --git a/docs/decommissioning/terraform.md b/docs/decommissioning/terraform.md new file mode 100644 index 000000000..e97677952 --- /dev/null +++ b/docs/decommissioning/terraform.md @@ -0,0 +1,13 @@ +# Decomissioning a DSS instance deployed with terraform + +## Clean up + +!!! danger + Note that the following operations can't be reverted and all data will be lost. + +1. Navigate to the workspace folder and run `./get-credentials.sh` to login to kubernetes; see corresponding [services deployment documentation](../deployment/services/after-terraform.md) +2. To delete all resources, run `tk delete .` in the workspace folder. +3. (AWS) Make sure that all [load balancers](https://eu-west-1.console.aws.amazon.com/ec2/home#LoadBalancers:) and [target groups](https://eu-west-1.console.aws.amazon.com/ec2/home#TargetGroups:) have been deleted from the AWS region before next step. +4. `terraform destroy` in your infrastructure folder. +5. (AWS) On the [EBS page](https://eu-west-1.console.aws.amazon.com/ec2/home#Volumes:), make sure to manually clean up the persistent storage. Note that the correct AWS region shall be selected. +6. (GCP) [Manually clean up the persistent storage](https://console.cloud.google.com/compute/disks) diff --git a/docs/deployment/.nav.yml b/docs/deployment/.nav.yml new file mode 100644 index 000000000..052079436 --- /dev/null +++ b/docs/deployment/.nav.yml @@ -0,0 +1,6 @@ +nav: + - "Overview": index.md + - "Infrastructure": infrastructure/ + - "Pooling": pooling/ + - "Services": services/ + - "Verification": verification/ diff --git a/docs/deployment/index.md b/docs/deployment/index.md new file mode 100644 index 000000000..9e14b8c0f --- /dev/null +++ b/docs/deployment/index.md @@ -0,0 +1,61 @@ +# Deployment of a DSS instance + +An operational DSS deployment requires a specific architecture to be compliant with [standards requirements](https://github.com/interuss/dss?tab=readme-ov-file#standards-and-regulations) and meet performance expectations as described in [architecture background](../background/architecture.md). +This section describes the deployment procedures recommended by InterUSS to achieve this compliance and meet these expectations. + +The deployment of a DSS instance involves four stages: + +1. Provisioning the required cloud resources, in particular a Kubernetes cluster: [**Infrastructure**](./infrastructure/index.md). + +1. Coordinating with other pool members (if any): [**Pooling**](./pooling/index.md). + +1. Deploying the DSS applications under the form of Kubernetes resources: [**Services**](./services/index.md). + +1. Verifying the deployment was completed successfully and is fit for purpose: [**Verification**](./verification/index.md). + +```mermaid +flowchart + Start --> Infrastructure --> Services --> Verification --> DeployedInstance + Start --> Pooling --> Services + + Start(Start) + Infrastructure["Infrastructure
(VMs, Kubernetes cluster, IP addresses, DNS resolution, etc)"] + Pooling["Pooling
(Coordinating with other DSS instances in the pool)"] + Services["Services
(Deployment of jobs to Kubernetes cluster)"] + Verification["Verification
(Verifying deployment was successful)"] + DeployedInstance["Deployed DSS instance"] + + click Infrastructure "./infrastructure" "Go to infrastructure overview" + click Pooling "./pooling" "Go to pooling overview" + click Services "./services" "Go to services overview" + click Verification "./verification" "Go to verification overview" +``` + +## Deployment Checklist + +This checklist outlines the major decisions and steps required to deploy a non-local (e.g. production, qualification) DSS instance. + +### Preparation + +* [ ] Decide on the datastore you will use (CockroachDB or YugabyteDB). **All participants in a DSS Pool must use the same datastore**, so plan accordingly. +* [ ] Decide how and where you will deploy [the infrastructure](./infrastructure/index.md) of your DSS instances: + * This repository provides Terraform configurations for [Amazon Web Services (EKS)](infrastructure/aws.md) and [Google Cloud (GKE)](infrastructure/google.md) to deploy a Kubernetes cluster (the infrastructure into which the Services will be deployed). +* [ ] Decide how and where you will deploy [the services](./services/index.md) of your DSS instances: + * This repository provides [Tanka](https://github.com/interuss/dss/blob/master/deploy/services/tanka/) files and [Helm Charts](https://github.com/interuss/dss/blob/master/deploy/services/helm-charts/dss) to be used to deploy Services into a Kubernetes cluster. Terraform will automatically generate these configurations if needed. +* [ ] Prepare sufficient resources for the services. + * In particular, review the [CockroachDB recommendations](https://www.cockroachlabs.com/docs/v24.1/recommended-production-settings#cloud-specific-recommendations) and [YugabyteDB recommendations](https://docs.yugabyte.com/stable/deploy/checklist/#public-clouds); the datastore will consume the majority of the resources. + * Example sizing is also describled in [sizing](../background/sizing.md). + +### Deployment + +* [ ] Deploy the [infrastructure](./infrastructure/index.md) by following the guides based on your previous infrastructure choice. +* [ ] Define the [pool](./pooling/index.md) in which your DSS will operate. +* [ ] Deploy the [services](./services/index.md) by following the guides based on your previous services choice. +* [ ] [Verify](./verification/index.md) the DSS instance was deployed successfully and is fit for purpose. + +### Operations + +Once a DSS instance is deployed, its owner will need to [operate it](../operations/index.md) appropriately to meet requirements. + +* [ ] Identify [operational procedures](../operations/index.md) that may be needed while operating the DSS instance. +* [ ] Ensure operating personnel are trained on all applicable operational procedures that may be needed. diff --git a/docs/infrastructure/.nav.yml b/docs/deployment/infrastructure/.nav.yml similarity index 53% rename from docs/infrastructure/.nav.yml rename to docs/deployment/infrastructure/.nav.yml index 76bd831ef..0dda1567b 100644 --- a/docs/infrastructure/.nav.yml +++ b/docs/deployment/infrastructure/.nav.yml @@ -1,7 +1,5 @@ -flatten_single_child_sections: true nav: - - "Introduction": index.md + - "Overview": index.md - "Amazon Web Services with terraform": aws.md - "Google Cloud Platform with terraform": google.md - - "Google Cloud Platform manually": google-manual.md - "Minikube": minikube.md diff --git a/docs/infrastructure/aws.md b/docs/deployment/infrastructure/aws.md similarity index 60% rename from docs/infrastructure/aws.md rename to docs/deployment/infrastructure/aws.md index 1c9744488..c9a002301 100644 --- a/docs/infrastructure/aws.md +++ b/docs/deployment/infrastructure/aws.md @@ -1,9 +1,8 @@ -# Deploy a DSS instance on Amazon Web Services to terraform +# Deploy a DSS instance on Amazon Web Services via terraform This terraform module creates a Kubernetes cluster in Amazon Web Services using the Elastic Kubernetes Service (EKS) and generates the tanka files to deploy a DSS instance. - ## Getting started ### Prerequisites @@ -36,7 +35,9 @@ Download & install the following tools to your workstation: 2. output.tf 3. terraform.tfvars 4. variables.gen.tf - 3. Set the variables in `terraform.tfvars` according to your environment. See [TFVARS.gen.md](https://github.com/interuss/dss/blob/master/deploy/infrastructure/modules/terraform-aws-dss/TFVARS.gen.md) for variables descriptions. + 3. Set the variables in `terraform.tfvars` according to your environment. + 1. See [TFVARS.gen.md](https://github.com/interuss/dss/blob/master/deploy/infrastructure/modules/terraform-aws-dss/TFVARS.gen.md) for variables descriptions. + 2. See [authentication documentation](../../background/authentication.md) for additional information. 4. Initialize terraform: `terraform init`. 5. Run `terraform plan` to check that the configuration is valid. It will display the resources which will be provisioned. 6. Run `terraform apply` to deploy the cluster. (This operation may take up to 15 min.) @@ -90,44 +91,6 @@ Download & install the following tools to your workstation: 3. Create the entries for SSL certificate validation according to the information provided in `gateway_address.certificate_validation_dns`. ---- - -## Deployment of the DSS services - -During the successful run, the terraform job has created a new [workspace](https://github.com/interuss/dss/tree/master/build/workspace) -for the cluster. The new workspace name corresponds to the cluster context. The cluster context -can be retrieved by running `terraform output` in your infrastructure folder (ie /deploy/infrastructure/personal/terraform-aws-dss-dev). - -It contains scripts to operate the cluster and setup the services. - -1. Go to the new workspace `/build/workspace/${cluster_context}`. - 1. Run `./get-credentials.sh` to login to kubernetes. You can now access the cluster with `kubectl`. - -2. Prepare the datastore certificates: - -=== "Yugabyte" - 1. Generate the certificates using `./dss-certs.sh init` - 1. If joining a cluster, check `dss-certs.sh`'s [help](../operations/certificates-management.md) to add others CA in your pool and share your CA with others pools members. - 1. Deploy the certificates using `./dss-certs.sh apply`. - -=== "CockroachDB" - 1. Generate the certificates using `./make-certs.sh`. Follow script instructions if you are not initializing the cluster. - 1. Deploy the certificates using `./apply-certs.sh`. - ---- -3. Go to the tanka workspace in `/deploy/services/tanka/workspace/${cluster_context}`. -4. Run `tk apply .` to deploy the services to kubernetes. (This may take up to 30 min) -5. Wait for services to initialize: - - On AWS, load balancers and certificates are created by Kubernetes Operators. Therefore, it may take few minutes (~5min) to get the services up and running and generate the certificate. To track this progress, go to the following pages and check that: - - On the [EKS page](https://eu-west-1.console.aws.amazon.com/eks/home), the status of the kubernetes cluster should be `Active`. - - On the [EC2 page](https://eu-west-1.console.aws.amazon.com/ec2/home#LoadBalancers:), the load balancers (1 for the gateway, 1 per cockroach nodes) are in the state `Active`. -6. Verify that basic services are functioning by navigating to https://your-gateway-domain.com/healthy. - - -## Clean up +## Next steps -1. Note that the following operations can't be reverted and all data will be lost. -2. To delete all resources, run `tk delete .` in the workspace folder. -3. Make sure that all [load balancers](https://eu-west-1.console.aws.amazon.com/ec2/home#LoadBalancers:) and [target groups](https://eu-west-1.console.aws.amazon.com/ec2/home#TargetGroups:) have been deleted from the AWS region before next step. -4. `terraform destroy` in your infrastructure folder. -5. On the [EBS page](https://eu-west-1.console.aws.amazon.com/ec2/home#Volumes:), make sure to manually clean up the persistent storage. Note that the correct AWS region shall be selected. +Proceed to [pooling configuration](../pooling/index.md). diff --git a/docs/infrastructure/google.md b/docs/deployment/infrastructure/google.md similarity index 60% rename from docs/infrastructure/google.md rename to docs/deployment/infrastructure/google.md index f928c4302..49de35400 100644 --- a/docs/infrastructure/google.md +++ b/docs/deployment/infrastructure/google.md @@ -46,7 +46,9 @@ This guide will help you deploy a DSS instance to Google Cloud Platform with ter 2. output.tf 3. terraform.tfvars 4. variables.gen.tf - 3. Set the variables in `terraform.tfvars` according to your environment. See [TFVARS.gen.md](https://github.com/interuss/dss/blob/master/deploy/infrastructure/modules/terraform-google-dss/TFVARS.gen.md) for variables descriptions. + 3. Set the variables in `terraform.tfvars` according to your environment. + 1. See [TFVARS.gen.md](https://github.com/interuss/dss/blob/master/deploy/infrastructure/modules/terraform-google-dss/TFVARS.gen.md) for variables descriptions. + 2. See [authentication documentation](../../background/authentication.md) for additional information. 4. Initialize terraform: `terraform init`. 5. Run `terraform plan` to check that the configuration is valid. It will display the resources which will be provisioned. 6. Run `terraform apply` to deploy the cluster. (This operation may take up to 15 min.) @@ -87,45 +89,6 @@ This guide will help you deploy a DSS instance to Google Cloud Platform with ter - `crdb_addresses[*].expected_dns` - `gateway_address.expected_dns` ---- +## Next steps -## Deployment of the DSS services - -Following the successful terraform run, you should find a new [workspace directory](https://github.com/interuss/dss/tree/master/build/workspace) -for the new cluster. The new workspace name corresponds to the cluster context. The cluster context -can be retrieved by running `terraform output` in your infrastructure folder (eg /deploy/infrastructure/personal/terraform-google-dss-dev). - -It contains scripts to operate the cluster and setup the services. - -1. Go to the new workspace `/build/workspace/${cluster_context}`. - 2. Run `./get-credentials.sh` to login to kubernetes. You can now access the cluster with `kubectl`. - -3. Prepare the datastore certificates: -=== "Yugabyte" - 1. Generate the certificates using `./dss-certs.sh init` - 1. If joining a cluster, check `dss-certs.sh`'s [help](../operations/certificates-management.md) to add others CA in your pool and share your CA with others pools members. - 1. Deploy the certificates using `./dss-certs.sh apply`. - -=== "CockroachDB" - 1. Generate the certificates using `./make-certs.sh`. Follow script instructions if you are not initializing the cluster. - 1. Deploy the certificates using `./apply-certs.sh`. - ---- - -5. Go to the tanka workspace in `/deploy/services/tanka/workspace/${cluster_context}`. -6. Run `tk apply .` to deploy the services to kubernetes. (This may take up to 30 min) -7. Wait for services to initialize: - - On Google Cloud, the highest-latency operation is provisioning of the HTTPS certificate which generally takes 10-45 minutes. To track this progress: - - Go to the "Services & Ingress" left-side tab from the Kubernetes Engine page. - - Click on the https-ingress item (filter by just the cluster of interest if you have multiple clusters in your project). - - Under the "Ingress" section for Details, click on the link corresponding with "Load balancer". - - Under Frontend for Details, the Certificate column for HTTPS protocol will have an icon next to it which will change to a green checkmark when provisioning is complete. - - Click on the certificate link to see provisioning progress. - - If everything indicates OK and you still receive a cipher mismatch error message when attempting to visit /healthy, wait an additional 5 minutes before attempting to troubleshoot further. -8. Verify that basic services are functioning by navigating to https://your-gateway-domain.com/healthy. - -## Clean up - -To delete all resources, run `terraform destroy`. Note that this operation can't be reverted and all data will be lost. - -For Google Cloud Engine, make sure to manually clean up the persistent storage: https://console.cloud.google.com/compute/disks +Proceed to [pooling configuration](../pooling/index.md). diff --git a/docs/deployment/infrastructure/index.md b/docs/deployment/infrastructure/index.md new file mode 100644 index 000000000..6bbf85ada --- /dev/null +++ b/docs/deployment/infrastructure/index.md @@ -0,0 +1,24 @@ +# Deploying DSS infrastructure + +This section describes how to deploy the infrastructure for a DSS instance. + +## Prerequisites + +Before beginning infrastructure deployment, download & install the following tools to your workstation: + +- [Install kubectl](https://kubernetes.io/docs/tasks/tools/install-kubectl/) to + interact with kubernetes + - Confirm successful installation with `kubectl version --client` (should + succeed from any working directory). + - Note that kubectl can alternatively be installed via the Google Cloud SDK + `gcloud` shell if using Google Cloud. + +## Deployment Options + +The DSS can be deployed on various platforms. Choose the method that best suits your needs: + +| Platform | Tools | Description | +| :--- | :--- | :--- | +| **Amazon Web Services** | Terraform | [Deploy on AWS using Terraform](aws.md) to provision EKS and required resources. | +| **Google Cloud Platform** | Terraform | [Deploy on GCP using Terraform](google.md) to provision GKE and required resources. | +| **Locally** | Minikube | [Deploy locally using Minikube](minikube.md) for development and testing. | diff --git a/docs/deployment/infrastructure/minikube.md b/docs/deployment/infrastructure/minikube.md new file mode 100644 index 000000000..571e9ca43 --- /dev/null +++ b/docs/deployment/infrastructure/minikube.md @@ -0,0 +1,34 @@ +# Deploy a DSS instance locally on Minikube + +This section provide instructions to prepare a local minikube cluster. + +Minikube is going to take care of most of the work by spawning a local kubernetes cluster. + +## Getting started + +### Prerequisites + +In addition to the [general infrastructure prerequisites](./index.md#prerequisites), download & install the following tools to your workstation: + +1. Install [minikube](https://minikube.sigs.k8s.io/docs/start/) (First step only). + +### Create a new minikube cluster + +1. Run `minikube start -p dss-local-cluster` to create a new cluster. +2. Run `minikube tunnel -p dss-local-cluster` and keep it running to expose LoadBalancer services. + +If needed, you can change the name of the cluster (`dss-local-cluster` in this documentation) as needed. You may also deploy multiple cluster at the same time, using different names. + +### Access to the cluster + +Minikube provide a UI, should you want to keep track of deployment and/or inspect the cluster. To start it, use the following command: + +1. `minikube dashboard -p dss-local-cluster` + +You can also use any other tool as needed. You can switch to the cluster's context by using the following command: + +1. `kubectl config use-context dss-local-cluster` + +## Next steps + +[Deploy services](../services/to-minikube.md) diff --git a/docs/deployment/pooling/.nav.yml b/docs/deployment/pooling/.nav.yml new file mode 100644 index 000000000..5fabf7d8a --- /dev/null +++ b/docs/deployment/pooling/.nav.yml @@ -0,0 +1,4 @@ +nav: + - "Overview": index.md + - "With CockroachDB": crdb.md + - "With YugabyteDB": yugabyte.md diff --git a/docs/deployment/pooling/crdb.md b/docs/deployment/pooling/crdb.md new file mode 100644 index 000000000..54b13e28f --- /dev/null +++ b/docs/deployment/pooling/crdb.md @@ -0,0 +1,99 @@ +# Pooling configuration with CockroachDB + +CockroachDB certificates define the boundaries of the pool and must be generated before services can be deployed. + +## Prerequisites + +* [Install CockroachDB 24.1](https://www.cockroachlabs.com/docs/v24.1/install-cockroachdb-linux) to + generate CockroachDB certificates. + - These instructions assume CockroachDB Core. + - You may need to run `sudo chmod +x /usr/local/bin/cockroach` after + completing the installation instructions. + - Confirm successful installation with `cockroach version` + +In the instructions below: + +* `$CLUSTER_CONTEXT` is the name of the Kubernetes cluster deployed as + [infrastructure](../infrastructure/index.md) onto which [services](../services/index.md) will be deployed. +* `$NAMESPACE` is the namespace for this DSS instance +* Each `ADDRESS` is the DNS entry for a CockroachDB node that will use the + certificates generated by the command. This is usually just the nodes + constituting this DSS instance, though if you maintain multiple DSS instances + in a single pool, the separate instances may share certificates. Note that + `--node-address` must include all the hostnames and/or IP addresses that + other CockroachDB nodes will use to connect to your nodes (the nodes using + these certificates). Wildcard notation is supported, so you can use + `*...com>`. If following the recommendations above, use a + single ADDRESS similar to `*.db.yourdeployment.yourdomain.com`. The ADDRESS + entries should be separated by spaces. + +## First or only DSS instance + +If this DSS instance will be the first and/or only DSS instance in the pool, follow the instructions in this section. + +Use [`make-certs.py` script](https://github.com/interuss/dss/blob/master/build/make-certs.py) to create certificates for + the CockroachDB nodes in this DSS instance: + + ./make-certs.py --cluster $CLUSTER_CONTEXT --namespace $NAMESPACE + --node-address
... + + 1. Note: If you are creating multiple DSS instances at once and joining + them together, you may want to copy the nth instance's `ca.crt` into + the rest of the instances, such that ca.crt is the same across all + instances. + +Proceed to "Apply certs" section below. + +## Joining an existing pool + +If this DSS instance is joining an existing pool, follow the instructions in this section. + +### Obtain existing deployment information + +Before certs for this DSS instance can be generated, the following information must be gathered from participants in the existing pool: + +* Public certificate of each DSS instance in the existing pool, concatenated into a single `ca.crt file` located at `CA_CERT_FILE` +* The address of each DSS instance in the existing pool + +### Generate certs + +Use [`make-certs.py` script](https://github.com/interuss/dss/blob/master/build/make-certs.py) to create certificates for + the CockroachDB nodes in this DSS instance: + + ./make-certs.py --cluster $CLUSTER_CONTEXT --namespace $NAMESPACE + --node-address
... + --ca-cert-to-join + +### Adopt certs + +Before services are deployed to this DSS instance, all existing participants in the pool must update their DSS instances to accept the certificates of this DSS instance (just generated). + +Share ca.crt with each participant in the existing pool and have them apply the new ca.crt, which now contains both your instance's and the original instances' public certs, to enable secure bi-directional communication. The operator of each existing DSS instance, upon receipt of the combined ca.crt from this joining instance, should perform the "Updating ca.crt" actions below and provide positive confirmation that they have completed those actions. + +!!! danger + Do not proceed until every operator of a DSS instance in the existing pool has positively confirmed that their DSS instance has adopted the provided ca.crt. + +#### Updating ca.crt + +When a participant seeking to deploy a DSS instance into an existing pool requests that a participant with an existing DSS instance in the pool adopt a provided ca.crt file, **the participant with the existing DSS instance should**: + +1. Overwrite its existing ca.crt with the new ca.crt provided by the DSS + instance joining the pool. +1. Upload the new ca.crt to its cluster using `./apply-certs.sh $CLUSTER_CONTEXT $NAMESPACE` (see top of page for `$CLUSTER_CONTEXT` and `$NAMESPACE`) +1. Restart their CockroachDB pods to recognize the updated ca.crt: + `kubectl rollout restart statefulset/cockroachdb --namespace $NAMESPACE` +1. Inform the prospective pool participant when the CockroachDB pods have finished restarting (typically around 10 minutes) + +## Apply certs + +Regardless of how certs were generated, apply/deploy them to the +[infrastructure](../infrastructure/index.md) using the +[`apply-certs.sh` script](https://github.com/interuss/dss/blob/master/build/apply-certs.sh) +to create secrets on the Kubernetes cluster containing the certificates and +keys generated in the previous step. + + ./apply-certs.sh $CLUSTER_CONTEXT $NAMESPACE + +## Next steps + +Proceed to [services deployment](../services/index.md). diff --git a/docs/deployment/pooling/index.md b/docs/deployment/pooling/index.md new file mode 100644 index 000000000..9373c20b0 --- /dev/null +++ b/docs/deployment/pooling/index.md @@ -0,0 +1,16 @@ +# Pooling for deployment + +Before [services](../services/index.md) can be deployed to the +[infrastructure](../infrastructure/index.md) of a DSS instance, the pool that the DSS +instance will create or join must be defined. This section describes how to +define that pooling configuration. See +[pooling background documentation](../../background/pooling.md) for a +conceptual overview. + +## Data stores + +The pooling procedure differs depending on data store used by the DSS +instance: + +- [CockroachDB pooling](./crdb.md) +- [YugabyteDB pooling](./yugabyte.md) diff --git a/docs/operations/certificates-management.md b/docs/deployment/pooling/yugabyte.md similarity index 51% rename from docs/operations/certificates-management.md rename to docs/deployment/pooling/yugabyte.md index 936aa7d0a..31ebdce8b 100644 --- a/docs/operations/certificates-management.md +++ b/docs/deployment/pooling/yugabyte.md @@ -1,6 +1,8 @@ -# Certificates management (Yugabyte) +# Pooling configuration with YugabyteDB -## Introduction +## Certificates management + +### Introduction The `dss-certs.py` helps you manage the set of certificates used for your DSS deployment. @@ -8,14 +10,14 @@ Should this DSS beeing part of a pool, the script also provide some helpers to m To run the script, just run `./dss-certs.py`. The python script don't require any dependencies, just a recent version of python 3. -## Quick start guide +### Quick start guide -### Single DSS instance in minikube` +#### Single DSS instance in minikube` * `./dss-certs.py --name test --cluster-context dss-local-cluster --namespace default init` * `./dss-certs.py --name test --cluster-context dss-local-cluster --namespace default apply` -### Pool of 3 DSS instances in minikube, in namespace `default`, `ns2` and `ns3` +#### Pool of 3 DSS instances in minikube, in namespace `default`, `ns2` and `ns3` * Creation of the 3 DSS instances certificates * `./dss-certs.py --name dss-instance-1 --cluster-context dss-local-cluster --namespace default init` @@ -34,51 +36,51 @@ To run the script, just run `./dss-certs.py`. The python script don't require an !!! Roll out restart required -## Operations +### Operations -### Common parameters +#### Common parameters -#### `--name` +##### `--name` The name of your DSS instance, that should identify it in a unique way. Used as main identifier for the set of certificates and in certificates. Example: `dss-west-1` -#### `--organization` +##### `--organization` The name of the organization managing the DSS Instance. Used in certificates generation. The combination of (name, organization) shall be unique in a cluster. Example: `Interuss` -#### `--cluster-context` +##### `--cluster-context` The kubernetes context the script should use. Example: `dss-local-cluster` -#### `--namespace` +##### `--namespace` The kubernetes namespace to use. Example: `default` -#### `--nodes-count` +##### `--nodes-count` The number of yugabyte nodes of your DSS instance. Default to `3`. -### `init` +#### `init` Initializes the certificates for a new DSS instance including a CA, a client certificate and a certificate for each yugabyte node. -### `apply` +#### `apply` Apply the current set of certificates to the kubernetes cluster. Shall be ran after each modification of the certificates, like addition / removal of CA in the pool, new `nodes-count` parameter. -### `regenerate-nodes` +#### `regenerate-nodes` Generate missing nodes certificates. Useful if you want to add new nodes in your DSS Instance. Don't forget to set the `nodes-count` parameters. -### `add-pool-ca` +#### `add-pool-ca` Add a CA certificate(s) of another(s) DSS Instance to the set of trusted certificates. Existing certificates are not added again. @@ -93,7 +95,7 @@ Examples: * `./dss-certs.py --name test --cluster-context dss-local-cluster --namespace default --ca-file /tmp/new-dss-ca add-pool-ca` * `./dss-certs.py --name test --cluster-context dss-local-cluster --namespace default get-pool-ca | ./dss-certs.py --name test2 --cluster-context dss-local-cluster --namespace namespace2 add-pool-ca` -### `remove-pool-ca` +#### `remove-pool-ca` Remove CA certificate(s) of DSS Instance(s) from the set of trusted certificates. Unknown certificates are not removed again. @@ -110,24 +112,118 @@ Example: * `./dss-certs.py --name test --cluster-context dss-local-cluster --namespace default remove-pool-ca --ca-serial="830ECFB0` * `./dss-certs.py --name test --cluster-context dss-local-cluster --namespace default remove-pool-ca --ca-serial="46548B7CC9699A7CFA54FF8FA85A619E830ECFB0` -### `list-pool-ca` +#### `list-pool-ca` List the set of accepted CA certificates. Also display a 'hash' of CA serial, that you may use to compare other DSS Instances list of CA certificates easily. -### `get-pool-ca` +#### `get-pool-ca` Return all CA certificate in the current pool. Can be used for debugging or to synchronize the set of CA certificates in a pool with others USS. -### `get-ca` +#### `get-ca` Return your own CA certificate . Display the compiled CA certificate. Can be used for debugging or to synchronize the set of CA certificates in a pool with others USS. -### `destroy` +#### `destroy` Destroy a certificate set. Be careful, there are no way to undo the command. + +## Deployment into a YugabyteDB pool + +### First instance + +Use [`dss-certs.py` script](#certificates-management) to create certificates for the Yugabyte nodes in this DSS instance. + +Each DSS instance must set `yugabyte_external_nodes` with the list of each +others DSS instance Yugabyte master nodes public endpoints, and CA certificates +must be exchanged. + +It's possible to have one DSS instance as starting point. In that case, +`yugabyte_external_nodes` will be empty and no CA exchange is needed. + +!!! info + Quick reminder for CA management: + + Each DSS instance should use `./dss-certs.sh init` To get the CA that should + be sent to others instances, use `./dss-certs.sh get-ca` To import the CA of + others DSS instance, use `./dss-certs.sh add-pool-ca` Finally, apply + certificates on the kubernetes cluster with `./dss-certs.sh apply` + +Ensure placement info is how you want it. See the section below for placement +requirements. + +### Joining an existing pool with a new instance + +They +will be joining an existing cluster, and they will need to request all CAs that +the pool is currently using (any one member of the pool may provide it). The +joining USS will also need a list of Yugabyte node addresses. + +The joining USS must create his own CA with `./dss-certs.sh init` and retrieve +it with `./dss-certs.sh get-ca`. This certificate must be provided to each +existing DSS instance in the pool that will import it with `./dss-certs.sh +add-pool-ca` and `./dss-certs.sh apply`. + +One of existing DSS instance shall provide to the joining USS all existing +certificate, using `./dss-certs.sh get-pool-ca`. The joining USS can import them +with `./dss-certs.sh add-pool-ca` and finally apply certificates with +`./dss-certs.sh apply`. As an alternative, each DSS instance can provide its +individual CA. + +Participants shall ensure they work with a coherent set of certificate by +comparing the pool CA hash. It is displayed after adding certificates or using +the `./dss-certs.sh list-pool-ca`. + +When CAs have been exchanged and configured everywhere, the joining participant +can bring his system online (e.g. by applying helm charts onto his cluster). The +`yugabyte_external_nodes` setting shall be set **before** starting the Yugabyte +cluster. + +New nodes shall be allowed into the cluster. For each new Yugabyte master node, +the following command shall be run on one master node of one existing DSS +instance : + +!!! warning + The `master_addresses` in all commands below must include the Yugabyte master + leader. Either always run commands in the cluster with the leader, or list all + public addresses. + +1. Connection to a master node: + + `kubectl exec -it yb-master-0 -- sh` + +1. Addition of a new master node + + ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 change_master_config ADD_SERVER [PUBLIC HOSTNAME] 7100`` + +The last command can be repeated as needed, however a small delay is needed for +the cluster to settle when adding a new node. If you get `Leader is not ready +for Config Change, can try again`, just try again. + +You should have all masters listed in the web ui or using the +``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 list_all_masters`` +command. + +The tserver nodes will join automatically, using the list of provided master +nodes. They can be listed for confirmation in the web ui or using the +``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 list_all_tablet_servers`` +command. + +The pool should then be re-verified for functionality +by running the prober test on each DSS instance, and the +[interoperability test scenario](https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/scenarios/astm/netrid/v19/dss_interoperability.md) +on the full pool (including the newly-added instance). + +Finally, the joining USS should provide its Yugabyte node addresses to all other +participants in the pool, and each other participant should add those addresses +to the `yugabyte_external_nodes` list their Yugabyte nodes will attempt to +contact upon restart. + +Ensure placement info is how you want it. See the section below for placement +requirements. diff --git a/docs/deployment/services/.nav.yml b/docs/deployment/services/.nav.yml new file mode 100644 index 000000000..d997349a2 --- /dev/null +++ b/docs/deployment/services/.nav.yml @@ -0,0 +1,5 @@ +nav: + - "Overview": index.md + - "After terraform": after-terraform.md + - "To Minikube": to-minikube.md + - "Support services": support.md diff --git a/docs/deployment/services/after-terraform.md b/docs/deployment/services/after-terraform.md new file mode 100644 index 000000000..e249fe912 --- /dev/null +++ b/docs/deployment/services/after-terraform.md @@ -0,0 +1,36 @@ +# Deploying services onto terraformed infrastructure + +Before following these instructions, ensure that a suitable docker image +[has been built and is available in an appropriate location](./index.md#docker-images). + +## Deployment of the DSS services + +Upon successful completion of [infrastructure deployment](../infrastructure/index.md) with terraform, a new +folder in [workspace](https://github.com/interuss/dss/tree/master/build/workspace) will have been +created for the cluster. The new workspace name corresponds to the cluster context. The cluster +context can be retrieved by running `terraform output` in your infrastructure folder (e.g., +/deploy/infrastructure/personal/terraform-aws-dss-dev or +/deploy/infrastructure/personal/terraform-google-dss-dev). + +It contains scripts to operate the cluster and setup the services. + +1. Go to the new workspace `/build/workspace/${cluster_context}`. +2. Run `./get-credentials.sh` to login to kubernetes. You can now access the cluster with `kubectl`. +3. Go to the tanka workspace in `/deploy/services/tanka/workspace/${cluster_context}`. +4. Run `tk apply .` to deploy the services to kubernetes. (This may take up to 30 min) +5. Wait for services to initialize: + - On AWS, load balancers and certificates are created by Kubernetes Operators. Therefore, it may take few minutes (~5min) to get the services up and running and generate the certificate. To track this progress, go to the following pages and check that: + - On the [EKS page](https://eu-west-1.console.aws.amazon.com/eks/home), the status of the kubernetes cluster should be `Active`. + - On the [EC2 page](https://eu-west-1.console.aws.amazon.com/ec2/home#LoadBalancers:), the load balancers (1 for the gateway, 1 per cockroach nodes) are in the state `Active`. + - On Google Cloud, the highest-latency operation is provisioning of the HTTPS certificate which generally takes 10-45 minutes. To track this progress: + - Go to the "Services & Ingress" left-side tab from the Kubernetes Engine page. + - Click on the https-ingress item (filter by just the cluster of interest if you have multiple clusters in your project). + - Under the "Ingress" section for Details, click on the link corresponding with "Load balancer". + - Under Frontend for Details, the Certificate column for HTTPS protocol will have an icon next to it which will change to a green checkmark when provisioning is complete. + - Click on the certificate link to see provisioning progress. + - If everything indicates OK and you still receive a cipher mismatch error message when attempting to visit /healthy, wait an additional 5 minutes before attempting to troubleshoot further. +6. Verify that basic services are functioning by navigating to https://your-gateway-domain.example.com/healthy. + +## Next steps + +Proceed to [support services](./support.md) to set up any permanent or recurring operations necessary to achieve requirements. diff --git a/docs/deployment/services/index.md b/docs/deployment/services/index.md new file mode 100644 index 000000000..4f52ddeed --- /dev/null +++ b/docs/deployment/services/index.md @@ -0,0 +1,95 @@ +# Deploying DSS services + +## Prerequisites + +Before beginning services deployment: + +- Deploy appropriate [infrastructure](../infrastructure/index.md) (Kubernetes cluster is available) +- Complete appropriate [pooling configuration](../pooling/index.md) +- Download & install the following tools to your workstation: + - [Install kubectl](https://kubernetes.io/docs/tasks/tools/install-kubectl/) to + interact with kubernetes + - Confirm successful installation with `kubectl version --client` (should + succeed from any working directory). + - Note that kubectl can alternatively be installed via the Google Cloud SDK + `gcloud` shell if using Google Cloud. + - [Install tanka](https://tanka.dev/install) + - On Linux, after downloading the binary per instructions, run + `sudo chmod +x /usr/local/bin/tk` + - Confirm successful installation with `tk --version` + - [Install Docker](https://docs.docker.com/get-docker/). + - Confirm successful installation with `docker --version` + +## Services deployment + +This section describes how to deploy services for a DSS instance once the +prerequisites have been satisified. + +Depending on [infrastructure deployment](../infrastructure/index.md) method, deploy services: + +- [After using terraform for infrastructure deployment](./after-terraform.md) +- [To a local Minikube cluster](./to-minikube.md) + +Then, deploy any necessary [support services](./support.md). + +## Docker images + +The application logic of the DSS is located in core-service which is provided in +a Docker image. To use the prebuilt InterUSS Docker images (without building +them yourself), simply use `docker.io/interuss/dss` for `VAR_DOCKER_IMAGE_NAME` +and the rest of this section may be skipped. + +Instead of using the prebuilt images, you can build the image locally and push +it to a Docker registry of your choice. All major cloud providers have a +docker registry service, or you can set up your own. + +To build these images locally and, optionally, push them to a docker registry: + +1. Set the environment variable `DOCKER_URL` to your docker registry url +endpoint. + + - For Google Cloud, `DOCKER_URL` should be set similarly to as described + [here](https://cloud.google.com/container-registry/docs/pushing-and-pulling#tag_the_local_image_with_the_registry_name), + like `gcr.io/your-project-id` (do not include the image name; + it will be appended by the build script) + + - For Amazon Web Services, `DOCKER_URL` should be set similarly to as described + [here](https://docs.aws.amazon.com/AmazonECR/latest/userguide/docker-push-ecr-image.html), + like `${aws_account_id}.dkr.ecr.${region}.amazonaws.com/` (do not include the image name; + it will be appended by the build script) + +1. Ensure you are logged into your docker registry service. + + - For Google Cloud, + [these](https://cloud.google.com/container-registry/docs/advanced-authentication#gcloud-helper) + are the recommended instructions (`gcloud auth configure-docker`). + Ensure that + [appropriate permissions are enabled](https://cloud.google.com/container-registry/docs/access-control). + + - For Amazon Web Services, create a private repository by following the instructions + [here](https://docs.aws.amazon.com/AmazonECR/latest/userguide/repository-create.html), then login + as described [here](https://docs.aws.amazon.com/AmazonECR/latest/userguide/docker-push-ecr-image.html). + +1. Use the [`build.sh` script](https://github.com/interuss/dss/blob/master/build/build.sh) in this directory to build and push + an image tagged with the current date and git commit hash. + +1. Note the VAR_* value printed at the end of the script. + +### Access to private repository + +See the description of `VAR_DOCKER_IMAGE_PULL_SECRET` in TFVARS.gen.md to configure authentication when using terraform to deploy [infrastructure](../infrastructure/index.md). + +### Verify signature of prebuilt InterUSS Docker images + +The prebuilt docker images are signed using [sigstore](https://www.sigstore.dev/). +The identity of the CI workflow, attested by GitHub, is used so sign the images. + +The signature may be verified by using [cosign](https://github.com/sigstore/cosign): +```shell +docker pull "docker.io/interuss/dss:latest" +cosign verify "docker.io/interuss/dss:latest" \ + --certificate-identity-regexp="https://github.com/interuss/dss/.github/workflows/dss-publish.yml@refs/*" \ + --certificate-oidc-issuer="https://token.actions.githubusercontent.com" +``` + +Adapt the version specified if required. diff --git a/docs/deployment/services/support.md b/docs/deployment/services/support.md new file mode 100644 index 000000000..46b640f79 --- /dev/null +++ b/docs/deployment/services/support.md @@ -0,0 +1,11 @@ +# Deploy support services + +In addition to the main services constituting a DSS instance, this page describes how to set up any permanent or recurring operations necessary to achieve requirements. + +* Review the [database cleanup documentation](../../operations/cleanup.md) and enable cleanup cron jobs if required. +* If needed, [monitor metrics](../../operations/monitoring.md) of your DSS instance. +* If needed, track the availability of your DSS instance using [health checks](../../operations/healthchecks.md). + +## Next steps + +Proceed to [verification](../verification/index.md) to ensure the new DSS deployment is working properly and fit for purpose. diff --git a/docs/deployment/services/to-minikube.md b/docs/deployment/services/to-minikube.md new file mode 100644 index 000000000..a9207f8bd --- /dev/null +++ b/docs/deployment/services/to-minikube.md @@ -0,0 +1,29 @@ +# Deployment of DSS services to Minikube + +## Upload or update local image + +Should you want to run the local docker image that you [built](./index.md#docker-images), run the following commands to upload / update your image + +1. `minikube image -p dss-local-cluster load interuss-local/dss` + +In the helm charts, use `docker.io/interuss-local/dss:latest` as image and be sure to set the `imagePullPolicy` to `Never`. + +## Deployment + +You can now deploy the DSS services using Helm or Tanka. See the repository `/deploy/services` for more information. + +=== "Helm" + Minikube specific settings: + + * Use the `global.cloudProvider` setting with the value `minikube` and deploy the charts on the `dss-local-cluster` kubernetes context. + +=== "Tanka" + An example configuration is provided in the repository: `/deploy/services/tanka/examples/minikube` + +--- + +To access the service, find the external IP using the `kubectl get services dss-dss-gateway` command. The port 80, without HTTPs is used. + +## Next steps + +[Decommission the Minikube instance](../../decommissioning/minikube.md) when it is no longer needed. diff --git a/docs/deployment/verification/.nav.yml b/docs/deployment/verification/.nav.yml new file mode 100644 index 000000000..3e3c5a206 --- /dev/null +++ b/docs/deployment/verification/.nav.yml @@ -0,0 +1,2 @@ +nav: + - "Overview": index.md diff --git a/docs/deployment/verification/index.md b/docs/deployment/verification/index.md new file mode 100644 index 000000000..aa559324c --- /dev/null +++ b/docs/deployment/verification/index.md @@ -0,0 +1,12 @@ +# Verification of successful DSS deployment + +Upon deployment completion, the following may be run against the DSS instance +to verify functionality: + + - The [USS qualifier](https://github.com/interuss/monitoring/tree/main/monitoring/uss_qualifier), + using a configuration based on the [DSS Probing dev configuration](https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/configurations/dev/dss_probing.yaml) + - The [prober test](https://github.com/interuss/monitoring/blob/main/monitoring/prober/README.md) + +## Next steps + +Proceed to [operations](../../operations/index.md) to prepare for any anticipated operational needs. diff --git a/docs/deployment_checklist.md b/docs/deployment_checklist.md deleted file mode 100644 index 86febf837..000000000 --- a/docs/deployment_checklist.md +++ /dev/null @@ -1,25 +0,0 @@ -# Deployment Checklist - -This checklist outlines the major decisions and steps required to deploy a non-local (e.g. production, qualification) DSS instance. - -## Preparation - -* [ ] Review the [architecture requirements](architecture/index.md). -* [ ] Decide on the datastore you will use (CockroachDB or YugabyteDB). **All participants in a DSS Pool must use the same datastore**, so plan accordingly. -* [ ] Decide how and where you will deploy your DSS instances: - * This repository provides Terraform configurations for [Amazon Web Services (EKS)](infrastructure/aws.md) and [Google Cloud (GKE)](infrastructure/google.md) to deploy a Kubernetes cluster (the infrastructure into which the Services will be deployed). - * This repository provides [Tanka](https://github.com/interuss/dss/blob/master/deploy/services/tanka/) files and [Helm Charts](https://github.com/interuss/dss/blob/master/deploy/services/helm-charts/dss) to be used to deploy Services into a Kubernetes cluster. Terraform will automatically generate these configurations if needed. - * You may also choose to deploy manually or use custom configuration tools. -* [ ] Prepare sufficient resources for the services. - * In particular, review the [CockroachDB recommendations](https://www.cockroachlabs.com/docs/v24.1/recommended-production-settings#cloud-specific-recommendations) and [YugabyteDB recommendations](https://docs.yugabyte.com/stable/deploy/checklist/#public-clouds); the datastore will consume the majority of the resources. - * Example sizing is also describled in [sizing](architecture/sizing.md). - - -## Deployment - -* [ ] Deploy the DSS instance by following the guides based on your previous infrastructure choices. - * [ ] If needed, you will pool your DSS instance. Guides are available for [CockroachDB](operations/pooling-crdb.md) and [YugabyteDB](operations/pooling.md). -* [ ] If needed, [monitor metrics](operations/monitoring.md) of your DSS instance. -* [ ] If needed, track the availability of your DSS instance using [health checks](operations/healthchecks.md). -* [ ] Review the [database cleanup documentation](operations/cleanup.md) and enable cleanup cron jobs if required. -* [ ] Review the [authentication documentation](operations/authentication.md) to configure how access tokens are verified. diff --git a/docs/documentation_design.md b/docs/documentation_design.md new file mode 100644 index 000000000..dcd5f470c --- /dev/null +++ b/docs/documentation_design.md @@ -0,0 +1,69 @@ +# Design of user documentation + +## Definitions + +The following definitions used: + +* A "DSS instance deployment" is a fully-commissioned, working instance of our DSS software, ready to fully meet all applicable product requirements + * So, for instance, `tk apply` doesn't necessarily achieve a completed DSS deployment if eviction needs to be set up before the product will meet all of its applicable requirements, like maintaining a reasonable level of performance over a long period of time. + * Since a DSS instance cannot exist without a pool and we do not support changing which pool a DSS instance uses after deployment, at least the initial pooling configuration is a component of deployment. +* "Infrastructure" is all of the cloud resources and configurations necessary to accept deployment of standard services (Kubernetes cluster, static IP addresses, DNS configurations, etc). +* "Services" is all of the executables necessary to produce a complete instance of our product when executed on a suitable infrastructure. +* "Operations" are actions performed on a DSS deployment that don't decommission that deployment. + * So, actions taken during the process of deploying a DSS instance are not "operations"; they are part of deployment. +* "Decommissioning" actions are performed on a DSS deployment to reduce the scope of a DSS deployment; especially to destroy that deployment entirely. +* "User" in this context is someone who wants to operate a DSS instance, but does not want to modify any part of the product (change software, change deployment tooling, etc). +* "User documentation" is the documentation relevant to the user defined above. + * Therefore, documentation only relevant to someone wanting to develop the software is not "user documentation". For instance, user documentation should probably not ask the user to install Go (but developer documentation likely would). +* "Background" is information useful to a user (as defined above) that is not part of procedural documentation. +* "Procedural documentation" is described in the section below. + +## User profile + +The user for all of the journeys below is assumed to be USS personnel with: +* A moderate level of general technical knowledge (sufficient to design, deploy, and operate the rest of their USS system apart from the DSS) +* A moderate level of familiarity with the relevant standard (e.g., ASTM F3411-22a), since the rest of their USS is implemented according to that standard +* A clear understanding of what the overall system needs to accomplish (e.g., scale and profile of load) +* Little to no knowledge of the InterUSS DSS implementation + * They will learn the information necessary to use our product from this documentation + +## Procedural documentation + +All of the user journeys below except background seek to accomplish a particular goal. To meet this need, documentation should present a clear and complete procedure for the user to follow. Starting at the top index of the user documentation, a user should be able to easily follow a linear path of actions to accomplish the goal, like a machine interpreter following a compiled binary. Procedural documentation is source code for human actions. + +The user can be easily directed to different documents (and sections of documents, to some extent) like a "goto" machine instruction. + +The user can be presented with "if blocks" as long as the user can unambiguously evaluate the condition, easily determine where to go under each condition, and easily see when the "if block" ends and where to go next. + +The user can be referred to "subroutines", but they must be provided with all information necessary to complete that "subroutine" ("arguments") before being referred, and the "subroutine" must have a clear endpoint that the user easily understands to mean that they return to where they were before entering the "subroutine". We should ideally limit the "stack depth" as much as practical, however, as keeping a large stack in working memory can be challenging for humans. + +At every step in the entire procedure, the user must have all the information necessary to execute that step before the step is introduced. If they are missing any information, an additional step must be added prior to that step to acquire the information. + +At every step in the entire procedure from the entrypoint to completion of the user journey, the next step the user should take must be clear. + +Noting to the user where they can learn more background information when they are interested is useful, but we should minimize the amount of mandatory work required to use our products. "Go learn how to use this proprietary tool well enough to translate your desired outcomes into tool usages" is highly undesirable; we should instead provide as close to the exact command the user should run using the proprietary tool as practical. + +Not all future user documentation needs to be procedural or background, but there should be a clear distinction between procedural and non-procedural documentation. For instance, "tool reference" describing parameters of command line tools and what they do would be useful non-procedural documentation if users might use those tools outside a procedure defined in procedural documentation, or if there are so many different outcomes different users may want to achieve that understanding the scope and capabilities of a tool is necessary for a user to select their desired outcome. However, if there isn't a user journey that involves knowing the details of tool usage, full documentation of the tool would likely be better suited to remain solely inside developer documentation. + +## User journeys + +### Top-level user journeys + +The documentation in this folder is built around a few primary user journeys: +* USS wants to deploy a DSS instance +* USS wants to operate an existing DSS instance +* USS wants to decommission an existing DSS instance +* USS wants to know more about the concepts behind parts of the product, why procedures are defined as they are, how they might deviate from InterUSS recommendations to fit their particular use case, etc + +These user journeys are fulfilled by the deployment, operations, decommissioning, and background folders. + +### Deployment user journey + +The deployment journey is split into four modules which are intended to be mostly separable. For instance, "services" describes how to start executables on a common infrastructure substrate, mostly regardless of how that infrastructure was created. There are some inter-module dependencies: for instance, service deployment via the tanka files automatically generated by terraform can only be used if terraform generated those files during infrastructure creation. + +* Infrastructure is the Kubernetes cluster and other cloud or local resources needed to run the services executables. +* Pooling is the set of steps necessary to configure how the services will establish or join the pool when they are first started. + * This is where the potentially-blocking async steps of coordinating with existing pool participants are located. +* Services is all the necessary executables, including support services like database eviction and monitoring. +* Verification is the check that the final product produced from all the previous modules was successfully deployed and is suitable for purpose. + * The configuration or other characteristics of the DSS deployment should not change after verification starts since doing so would mean the verification did not verify the final form of the deployment. diff --git a/docs/index.md b/docs/index.md index f4a57b14d..30dc4470e 100644 --- a/docs/index.md +++ b/docs/index.md @@ -1,45 +1,12 @@ -# DSS Deployment User Documentation +# DSS User Documentation ## Introduction -This website provides instructions to deploy the InterUSS USS to USS Discovery and Synchronization service. +This website provides instructions to use the InterUSS USS to USS Discovery and Synchronization product. -An operational DSS deployment requires a specific architecture to be compliant with [standards requirements](https://github.com/interuss/dss?tab=readme-ov-file#standards-and-regulations) and meet performance expectations as described in [architecture](architecture/index.md). -This page describes the deployment procedures recommended by InterUSS to achieve this compliance and meet these expectations. +## Sections - -## Getting started - -- Review [architecture requirements](architecture/index.md) -- Deploy a DSS instance to [Amazon Web Services (EKS)](infrastructure/aws.md) using terraform -- Deploy a DSS instance to [Google (GKE)](infrastructure/google.md) using terraform -- Deploy a DSS instance to [Google (GKE)](infrastructure/google-manual.md) manually step by step -- Deploy a DSS instance to [Minikube](infrastructure/minikube.md) - -## Tooling - -The deployment of a DSS instance involves 3 stages: - -1. Provisioning the required cloud resources, in particular a Kubernetes cluster: **The Infrastructure**. - -1. Deploying the DSS applications under the form of kubernetes resources: **The Services**. - -1. Recommending procedures and guidelines on how to operate the DSS: **The Operations**. - -![Deployment layers](assets/generated/deployment_layers.png) - -Depending on your level of expertise and your internal organizational practices, you should be able to use each layer independently or complementary. - -InterUSS offers two terraform modules to deploy the **Infrastructure**: - -- [Amazon Web Services](https://github.com/interuss/dss/blob/master/deploy/infrastructure/modules/terraform-aws-dss/) -- [Google Cloud Platform](https://github.com/interuss/dss/blob/master/deploy/infrastructure/modules/terraform-google-dss/) - -The **Services** are deployed using the following tools: - -- [Tanka](https://github.com/interuss/dss/blob/master/deploy/services/tanka/) -- [Helm Chart](https://github.com/interuss/dss/blob/master/deploy/services/helm-charts/dss) - -See [Operate a DSS instance](operations/index.md) for more information on tools to perform the **Operations**. - -A [deployment check list](deployment_checklist.md) is also available to help you deploy your first instance. +- [Deployment](./deployment/index.md): Instructions to deploy a new DSS instance +- [Operations](./operations/index.md): Recommended procedures and guidelines on how to operate a deployed DSS instance +- [Decommissioning](./decommissioning/index.md): Instructions to decommission a deployed DSS instance +- [Background](./background/index.md): More detailed background information to understand how the product works diff --git a/docs/infrastructure/google-manual.md b/docs/infrastructure/google-manual.md deleted file mode 100644 index cb69d55c2..000000000 --- a/docs/infrastructure/google-manual.md +++ /dev/null @@ -1,490 +0,0 @@ -# Deploying a DSS instance to Google Cloud Platform manually step by step - -This document describes how to deploy a production-style DSS instance to -interoperate with other DSS instances in a DSS pool. - -## Preface - -This doc describes a procedure for deploying the DSS and its dependencies -(namely CockroachDB) via Kubernetes. The use of Kubernetes is not a requirement, -and a DSS instance can join a CRDB cluster constituting a DSS pool as long as it -meets the CockroachDB requirements below. - -## Prerequisites - -Install the required tools as described [here](./index.md#prerequisites). - -## Deploying a DSS instance via Kubernetes - -It discusses deploying a Kubernetes service manually, although you can deploy -a DSS instance however you like as long as it meets the CockroachDB requirements -above. You can do this on any supported -[cloud provider](https://kubernetes.io/docs/concepts/cluster-administration/cloud-providers/) -or even on your own infrastructure. Consult the Kubernetes documentation for -your chosen provider. - -1. Create a new Kubernetes cluster. We recommend a new cluster for each DSS - instance. A reasonable cluster name might be `dss-us-prod-e4a` (where `e4a` - is a zone identifier abbreviation), `dss-ca-staging`, - `dss-mx-integration-sae1a`, etc. The name of this cluster will be combined - with other information by Kubernetes to generate a longer cluster context - ID. - - - On Google Cloud, the recommended procedure to create a cluster is: - - In Google Cloud Platform, go to the Kubernetes Engine page and under - Clusters click Create cluster. - - Name the cluster appropriately; e.g., `dss-us-prod` - - Select Zonal and [a compute-zone appropriate to your - geography](https://cloud.google.com/compute/docs/regions-zones#available) - - For the "default-pool" node pool: - - Enter 3 for number of nodes. - - In the "Nodes" bullet under "default-pool", select N2 series and - n2-standard-4 for machine type. - - In the "Networking" bullet under "Clusters", ensure "Enable [VPC - -native traffic](https://cloud.google.com/kubernetes-engine/docs/how-to/alias-ips)" - is checked. - -1. Make sure correct cluster context is selected by printing the context - name to the console: `kubectl config current-context` - - - Record this value and use it for `$CLUSTER_CONTEXT` below; perhaps: - `export CLUSTER_CONTEXT=$(kubectl config current-context)` - - - On Google Cloud, first configure kubectl to interact with the cluster - created above with - [these instructions](https://cloud.google.com/kubernetes-engine/docs/quickstart). - Specifically: - - `gcloud config set project your-project-id` - - `gcloud config set compute/zone your-compute-zone` - - `gcloud container clusters get-credentials your-cluster-name` - -1. Ensure the desired namespace is selected; the recommended - namespace is simply `default` with one cluster per DSS instance. Print the - the current namespaces with `kubectl get namespace`. Use the current - namespace as the value for `$NAMESPACE` below; perhaps use an environment - variable for convenience: `export NAMESPACE=`. - - It may be useful to create a `login.sh` file with content like that shown - below and `source login.sh` when working with this cluster. - - GCP: - ```shell - #!/bin/bash - - export CLUSTER_NAME= - export REGION= - gcloud config set project - gcloud config set compute/zone $REGION-a - gcloud container clusters get-credentials $CLUSTER_NAME - export CLUSTER_CONTEXT=$(kubectl config current-context) - export NAMESPACE=default - export DOCKER_URL=docker.io/interuss - echo "Current CLUSTER_CONTEXT is $CLUSTER_CONTEXT - ``` - -1. Create static IP addresses: one for the Core Service ingress, and one - for each CockroachDB node if you want to be able to interact with other - DSS instances. - - - If using Google Cloud, the Core Service ingress needs to be created as - a "Global" IP address, but the CRDB ingresses as "Regional" IP addresses. - IPv4 is recommended as IPv6 has not yet been tested. Follow - [these instructions](https://cloud.google.com/compute/docs/ip-addresses/reserve-static-external-ip-address#reserve_new_static) - to reserve the static IP addresses. Specifically (replacing - CLUSTER_NAME as appropriate since static IP addresses are defined at - the project level rather than the cluster level), e.g.: - - - `gcloud compute addresses create ${CLUSTER_NAME}-backend --global --ip-version IPV4` - - `gcloud compute addresses create ${CLUSTER_NAME}-crdb-0 --region $REGION` - - `gcloud compute addresses create ${CLUSTER_NAME}-crdb-1 --region $REGION` - - `gcloud compute addresses create ${CLUSTER_NAME}-crdb-2 --region $REGION` - -1. Link static IP addresses to DNS entries. - - - Your CockroachDB nodes should have a common hostname suffix; e.g., - `*.db.interuss.com`. Recommended naming is - `0.db.yourdeployment.yourdomain.com`, - `1.db.yourdeployment.yourdomain.com`, etc. - - - If using Google Cloud, see - [these instructions](https://cloud.google.com/dns/docs/quickstart#create_a_new_record) - to create DNS entries for the static IP addresses created above. To list - the IP addresses, use `gcloud compute addresses list`. - -1. [](){ #certificates }(Only if you use CockroachDB) Use [`make-certs.py` script](https://github.com/interuss/dss/blob/master/build/make-certs.py) to create certificates for - the CockroachDB nodes in this DSS instance: - - ./make-certs.py --cluster $CLUSTER_CONTEXT --namespace $NAMESPACE - [--node-address
...] - [--ca-cert-to-join ] - - 1. `$CLUSTER_CONTEXT` is the name of the cluster (see step 2 above). - - 1. `$NAMESPACE` is the namespace for this DSS instance (see step 3 above). - - 1. `Each ADDRESS` is the DNS entry for a CockroachDB node that will use the - certificates generated by this command. This is usually just the nodes - constituting this DSS instance, though if you maintain multiple DSS - instances in a single pool, the separate instances may share - certificates. Note that `--node-address` must include all the hostnames - and/or IP addresses that other CockroachDB nodes will use to connect to - your nodes (the nodes using these certificates). Wildcard notation is - supported, so you can use `*...com>`. If following - the recommendations above, use a single ADDRESS similar to - `*.db.yourdeployment.yourdomain.com`. The ADDRESS entries should be - separated by spaces. - - 1. If you are pooling with existing DSS instance(s) you need their CA - public cert (ca.crt), which will be concatenated with yours. Set - `--ca-cert-to-join` to a `ca.crt` file. Reach out to existing operators - to request their public cert. If not joining an existing pool, omit - this argument. - - 1. Note: If you are creating multiple DSS instances at once, and joining - them together you likely want to copy the nth instance's `ca.crt` into - the rest of the instances, such that ca.crt is the same across all - instances. - -1. (Only if you use Yugabyte) Use [`dss-certs.py` script](../operations/certificates-management.md) to create certificates for the Yugabyte nodes in this DSS instance. - -1. If joining an existing DSS pool, share ca.crt with the DSS instance(s) you - are trying to join, and have them apply the new ca.crt, which now contains - both your instance's and the original instance's public certs, to enable - secure bi-directional communication. Each original DSS instance, upon - receipt of the combined ca.crt from the joining instance, should perform the - actions below. While they are performing those actions, you may continue - with the instructions. - - 1. If you use CockroachDB: - - 1. Overwrite its existing ca.crt with the new ca.crt provided by the DSS - instance joining the pool. - 1. Upload the new ca.crt to its cluster using - `./apply-certs.sh $CLUSTER_CONTEXT $NAMESPACE` - 1. Restart their CockroachDB pods to recognize the updated ca.crt: - `kubectl rollout restart statefulset/cockroachdb --namespace $NAMESPACE` - 1. Inform you when their CockroachDB pods have finished restarting - (typically around 10 minutes) - - 1. If you use Yugabyte - - 1. Share your CA with `./dss-certs.py get-ca` - 1. Add others CAs of the pool with `./dss-certs.py add-pool-ca` - 1. Upload the new CAs to its cluster using - `./dss-certs.py apply` - 1. Restart their Yugabyte pods to recognize the updated ca.crt: - `kubectl rollout restart statefulset/yb-master --namespace $NAMESPACE` - `kubectl rollout restart statefulset/yb-tserver --namespace $NAMESPACE` - 1. Inform you when their Yugabyte pods have finished restarting - (typically around 10 minutes) - -1. Ensure the Docker images are built according to the instructions in the - [getting started](index.md#docker-images). - -1. From this working directory, - `cp -r ../deploy/services/tanka/examples/minimum/* workspace/$CLUSTER_CONTEXT`. Note that - the `workspace/$CLUSTER_CONTEXT` folder should have already been created - by the `make-certs.py` script. - Replace the imports at the top of `main.jsonnet` to correctly locate the files: - ``` - local dss = import '../../../deploy/services/tanka/dss.libsonnet'; - local metadataBase = import '../../../deploy/services/tanka/metadata_base.libsonnet'; - ``` - -1. If providing a .pem file directly as the public key to validate incoming - access tokens, copy it to [dss/build/jwt-public-certs](https://github.com/interuss/dss/tree/master/build/jwt-public-certs). - Public key specification by JWKS is preferred; if using the JWKS approach - to specify the public key, skip this step. - -1. Edit `workspace/$CLUSTER_CONTEXT/main.jsonnet` and replace all `VAR_*` - instances with appropriate values: - - 1. `VAR_NAMESPACE`: Same `$NAMESPACE` used in the make-certs.py (and - apply-certs.sh) scripts. - - 1. `VAR_CLUSTER_CONTEXT`: Same $CLUSTER_CONTEXT used in the `make-certs.py` - and `apply-certs.sh` scripts. - - 1. `VAR_ENABLE_SCD`: Set this boolean true to enable strategic conflict - detection functionality (currently an R&D project tracking an initial - draft of the upcoming ASTM standard). - - 1. `VAR_ENABLE_SCD_GLOBAL_LOCK`: Set this boolean true to enable - experimental global lock when working with SCD subscriptions. Reduc - e global throughput but improve throughput with lot of subscriptions in - the same areas. - - 1. `VAR_ENABLE_TIME_BASED_NOTIFICATION_INDEX`: Set this boolean true to enable - time-based notification index when working with RID and SCD subscriptions. - - 1. `VAR_ENABLE_DSS_METRICS`: Set this boolean true to enable - prometheus-compatible metric endpoint. - - 1. `VAR_LOCALITY`: Unique name for your DSS instance. Currently, we - recommend "_", and the `=` character is not - allowed. However, any unique (among all other participating DSS - instances) value is acceptable. - - 1. `VAR_DB_HOSTNAME_SUFFIX`: The domain name suffix shared by all of your - CockroachDB nodes. For instance, if your CRDB nodes were addressable at - `0.db.example.com`, `1.db.example.com`, and `2.db.example.com`, then - VAR_DB_HOSTNAME_SUFFIX would be `db.example.com`. - - 1. `VAR_DATASTORE`: Datastore to use. Can be set to 'cockroachdb' or 'yugabyte'. - - 1. `VAR_DATASTORE_MAX_OPEN_CONNS`: Maximum number of open connections to the database, default is 4. - - 1. `VAR_CRDB_DOCKER_IMAGE_NAME`: Docker image of cockroach db pods. Until - DSS v0.16, the recommended CockroachDB image name is `cockroachdb/cockroach:v21.2.7`. - From DSS v0.17, the recommended CockroachDB version is `cockroachdb/cockroach:v24.1.3`. - - 1. `VAR_CRDB_NODE_IPn`: IP address (**numeric**) of nth CRDB node (add more - entries if you have more than 3 CRDB nodes). Example: `1.1.1.1` - - 1. `VAR_SHOULD_INIT`: Set to `false` if joining an existing pool, `true` - if creating the first DSS instance for a pool. When set `true`, this - can initialize the data directories on your cluster, and prevent you - from joining an existing pool. - - 1. `VAR_EXTERNAL_CRDB_NODEn`: Fully-qualified domain name of existing CRDB - nodes if you are joining an existing pool. If more than three are - available, add additional entries. If not joining an existing pool, - comment out this `JoinExisting:` line. - - - You should supply a minimum of 3 seed nodes to every CockroachDB node. - These 3 nodes should be the same for every node (ie: every node points - to node 0, 1, and 2). For external DSS instances you should point to a - minimum of 3, or you can use a loadbalanced hostname or IP address of - other DSS instances. You should do this for every DSS instance in the - pool, including newly joined instances. See CockroachDB's note on the - [join flag](https://www.cockroachlabs.com/docs/stable/start-a-node.html#flags). - - 1. `VAR_YUGABYTE_DOCKER_IMAGE_NAME`: Docker image of Yugabyte db pods. - Ensure you use an image with mtls working, see https://github.com/yugabyte/yugabyte-db/commit/89685fa888daca54eb3164a8c301e3bda8cf41b0 - `interuss/yugabyte:2025.1.2.1-interuss` is a known good image. - - 1. `VAR_YUGABYTE_MASTER_IPn`: IP address (**numeric**) of nth Yugabyte - master node (add more entries if you have more than 3 nodes). - Example: `1.1.1.1` - - 1. `VAR_YUGABYTE_TSERVER_IPn`: IP address (**numeric**) of nth Yugabyte - tserver node (add more entries if you have more than 3 nodes). - Example: `1.1.1.1` - - 1. `VAR_YUGABYTE_MASTER_ADDRESSn`: List of addresses of Yugabyte master - nodes in the DSS pool. Must be accessible from all master/tserver nodes - and identical in a cluster. Example: `["0.master.db.uss1.example.com", "1.master.db.uss1.example.com", "3.master.db.uss1.example.com", "0.master.db.uss2.example.com", "1.master.db.uss2.example.com", "3.master.db.uss2.example.com"]` - You may remove this setting if you only have a simple 3-nodes local cluster. - - 1. `VAR_YUGABYTE_MASTER_RPC_BIND_ADDRESSES`: Bind address for yugabyte - master node. May use `${HOSTNAME}`, `${NAMESPACE}` or `${HOSTNAMENO}` - to use respectively hostname, namespace or number of the node. - Example: `${HOSTNAMENO}.master.db.uss1.example.com` - You may remove this setting if you only have a simple 3-nodes local cluster. - - 1. `VAR_YUGABYTE_MASTER_BROADCAST_ADDRESSES`: Broadcast address for yugabyte - master node. May use `${HOSTNAME}`, `${NAMESPACE}` or `${HOSTNAMENO}` - to use respectively hostname, namespace or number of the node. - Example: `${HOSTNAMENO}.master.db.uss1.example.com:7100` - You may remove this setting if you only have a simple 3-nodes local cluster. - - 1. `VAR_YUGABYTE_TSERVER_RPC_BIND_ADDRESSES`: Bind address for yugabyte - tserver node. May use `${HOSTNAME}`, `${NAMESPACE}` or `${HOSTNAMENO}` - to use respectively hostname, namespace or number of the node. - Example: `${HOSTNAMENO}.tserver.db.uss1.example.com` - You may remove this setting if you only have a simple 3-nodes local cluster. - - 1. `VAR_YUGABYTE_TSERVER_BROADCAST_ADDRESSES`: Broadcast address for yugabyte - tserver node. May use `${HOSTNAME}`, `${NAMESPACE}` or `${HOSTNAMENO}` - to use respectively hostname, namespace or number of the node. - Example: `${HOSTNAMENO}.tserver.db.uss1.example.com:9100` - You may remove this setting if you only have a simple 3-nodes local cluster. - - 1. `VAR_YUGABYTE_FIX_27367_ISSUE`: Fix issue [27367](https://github.com/yugabyte/yugabyte-db/issues/27367) - To make the fix working, RPC bind and broadcast addresses must be set to - the same, public value on where the master / tserver node is accessible. - - 1. `VAR_YUGABYTE_LIGHT_RESOURCES`: Use light resources in term of CPU/Memory - for Yugabyte nodes. You may use that for development purposes, to deploy - a Yugabyte in a small cluster to save costs and resources. - - 1. `VAR_YUGABYTE_PLACEMENT_CLOUD`: Yugabyte placement's cloud value, for - master and tserver nodes. - Example: `cloud-1` - - 1. `VAR_YUGABYTE_PLACEMENT_REGION`: Yugabyte placement's region value, for - master and tserver nodes. - Example: `uss-1` - - 1. `VAR_YUGABYTE_PLACEMENT_ZONE`: Yugabyte placement's zone value, for - master and tserver nodes. - Example: `zone-1` - - 1. `VAR_STORAGE_CLASS`: Kubernetes Storage Class to use for CockroachDB, - Yugabyte and Prometheus volumes. You can check your cluster's possible - values with `kubectl get storageclass`. If you're not sure, each cloud - provider has some default storage classes that should work: - - Google Cloud: `standard` - - Azure: `default` - - AWS: `gp3` - - 1. `VAR_INGRESS_NAME`: If using Google Kubernetes Engine, set this to the - the name of the core-service static IP address created above (e.g., - `CLUSTER_NAME-backend`). - - 1. `VAR_DOCKER_IMAGE_NAME`: Full name of the docker image built in the - section above. `build.sh` prints this name as the last thing it does - when run with `DOCKER_URL` set. It should look something like - `gcr.io/your-project-id/dss:2020-07-01-46cae72cf` if you built the image - yourself, or `docker.io/interuss/dss` if using the InterUSS image - without `build.sh`. - - - Note that `VAR_DOCKER_IMAGE_NAME` is used in two places. - - 1. `VAR_DOCKER_IMAGE_PULL_SECRET`: Secret name of the credentials to access - the image registry. If the image specified in VAR_DOCKER_IMAGE_NAME does not require - authentication to be pulled, then do not populate this instance and do not uncomment - the line containing it. You can use the following command to store the credentials - as kubernetes secret: - - > kubectl create secret -n VAR_NAMESPACE docker-registry VAR_DOCKER_IMAGE_PULL_SECRET \ - --docker-server=DOCKER_REGISTRY_SERVER \ - --docker-username=DOCKER_USER \ - --docker-password=DOCKER_PASSWORD \ - --docker-email=DOCKER_EMAIL - - For docker hub private repository, use `docker.io` as `DOCKER_REGISTRY_SERVER` and an - [access token](https://hub.docker.com/settings/security) as `DOCKER_PASSWORD`. - - 1. `VAR_APP_HOSTNAME`: Fully-qualified domain name of your Core Service - ingress endpoint. For example, `dss.example.com`. - - 1. `VAR_PUBLIC_ENDPOINT`: URL to publicly access your Core Service - ingress endpoint. For example, `https://dss.example.com`. Only for versions >=0.21. - - 1. `VAR_PUBLIC_KEY_PEM_PATH`: If providing a .pem file directly as the - public key to validate incoming access tokens, specify the name of this - .pem file here as `/jwt-public-certs/YOUR-KEY-NAME.pem` replacing - YOUR-KEY-NAME as appropriate. For instance, if using the provided - [`us-demo.pem`](https://github.com/interuss/dss/tree/master/build/jwt-public-certs/us-demo.pem), use the path - `/jwt-public-certs/us-demo.pem`. Note that your .pem file must have - been copied into [`jwt-public-certs`](https://github.com/interuss/dss/tree/master/build/jwt-public-certs) in an earlier - step, or mounted at runtime using a volume. - - - If providing an access token public key via JWKS, provide a blank - string for this parameter. - - 1. `VAR_JWKS_ENDPOINT`: If providing the access token public key via JWKS, - specify the JWKS endpoint here. Example: - `https://auth.example.com/.well-known/jwks.json` - - - If providing a .pem file directly as the public key to valid incoming access tokens, provide a blank string for this parameter. - - 1. `VAR_JWKS_KEY_ID`: If providing the access token public key via JWKS, - specify the `kid` (key ID) of they appropriate key in the JWKS file - referenced above. - - - If providing a .pem file directly as the public key to valid incoming access tokens, provide a blank string for this parameter. - - - If you are only turning up a single DSS instance for development, you - may optionally change `single_cluster` to `true`. - - 1. `VAR_SSL_POLICY`: When deploying on Google Cloud, a [ssl policy](https://cloud.google.com/load-balancing/docs/ssl-policies-concepts) - can be applied to the DSS Ingress. This can be used to secure the TLS connection. - Follow the [instructions](https://cloud.google.com/load-balancing/docs/use-ssl-policies) to create the Global SSL Policy and - replace VAR_SSL_POLICY variable with its name. `RESTRICTED` profile is recommended. - Leave it empty if not applicable. - - 1. `VAR_ENABLE_SCHEMA_MANAGER`: Set this to true to enable the schema manager jobs. - It is required to perform schema upgrades. Note that it is automatically enabled when `VAR_SHOULD_INIT` is true. - - 1. `VAR_EVICT_ENABLE_SCD_CRON`: Set this to true to enable the cron job that automatically cleanup expired SCD entries. - - 1. `VAR_EVICT_SCD_SCHEDULE`: When the SCD cleanup job shall be performed; expressed in cron format (https://crontab.guru/). - - 1. `VAR_EVICT_SCD_TTL`: How long expired SCD items should stay before being automatically removed; expressed in Go duration format (https://pkg.go.dev/time#ParseDuration). - - 1. `VAR_EVICT_SCD_ENABLE_OPERATIONAL_INTENTS`: Set this to true to enable cleanup of SCD operational intents. - - 1. `VAR_EVICT_SCD_ENABLE_SUBSCRIPTIONS`: Set this to true to enable cleanup of SCD subscriptions. - - 1. `VAR_EVICT_ENABLE_RID_CRON`: Set this to true to enable the cron job that automatically cleanup RID entries. - - 1. `VAR_EVICT_RID_SCHEDULE`: When the RID cleanup job shall be performed; expressed in cron format (https://crontab.guru/). - - 1. `VAR_EVICT_RID_TTL`: How long expired RID items should stay before being automatically removed; expressed in Go duration format (https://pkg.go.dev/time#ParseDuration). - - 1. `VAR_EVICT_RID_ENABLE_ISAS`: Set this to true to enable cleanup of RID ISAs. - - 1. `VAR_EVICT_RID_ENABLE_SUBSCRIPTIONS`: Set this to true to enable cleanup of RID subscriptions. - - -1. Edit workspace/$CLUSTER_CONTEXT/spec.json and replace all VAR_* - instances with appropriate values: - - 1. VAR_API_SERVER: Determine this value with the command: - - `echo $(kubectl config view -o jsonpath="{.clusters[?(@.name==\"$CLUSTER_CONTEXT\")].cluster.server}")` - - - Note that `$CLUSTER_CONTEXT` should be replaced with your actual - `CLUSTER_CONTEXT` value prior to executing the above command if you - have not defined a `CLUSTER_CONTEXT` environment variable. - - 1. VAR_NAMESPACE: See previous section. - -1. Use the [`apply-certs.sh` script](https://github.com/interuss/dss/blob/master/build/apply-certs.sh) to create secrets on the - Kubernetes cluster containing the certificates and keys generated in the - previous step. - - ./apply-certs.sh $CLUSTER_CONTEXT $NAMESPACE - -1. Run `tk apply workspace/$CLUSTER_CONTEXT` to apply it to the - cluster. - - - If you are joining an existing pool, do not execute this command until the - the existing DSS instances all confirm that their CockroachDB pods have - finished their rolling restarts. - -1. Wait for services to initialize. Verify that basic services are functioning - by navigating to https://your-domain.example.com/healthy. - - - On Google Cloud, the highest-latency operation is provisioning of the - HTTPS certificate which generally takes 10-45 minutes. To track this - progress: - - Go to the "Services & Ingress" left-side tab from the Kubernetes Engine - page. - - Click on the `https-ingress` item (filter by just the cluster of - interest if you have multiple clusters in your project). - - Under the "Ingress" section for Details, click on the link corresponding - with "Load balancer". - - Under Frontend for Details, the Certificate column for HTTPS protocol - will have an icon next to it which will change to a green checkmark when - provisioning is complete. - - Click on the certificate link to see provisioning progress. - - If everything indicates OK and you still receive a cipher mismatch error - message when attempting to visit /healthy, wait an additional 5 minutes - before attempting to troubleshoot further. - -1. If joining an existing pool, share your CRDB node addresses with the - operators of the existing DSS instances. They will add these node addresses - to JoinExisting where `VAR_CRDB_EXTERNAL_NODEn` is indicated in the minimum - example, and then update their deployment: - - `tk apply workspace/$CLUSTER_CONTEXT` - -## Pooling - -See [the pooling documentation](../operations/pooling.md). - -## Tools - -See [operations monitoring documentation](../operations/monitoring.md). - -## Troubleshooting - -See [Troubleshooting in `deploy/operations`](../operations/troubleshooting.md). - -### Garbage collector job -Only since commit [c789b2b](https://github.com/interuss/dss/commit/c789b2b4a9fa5fb651d202da0a3abc02a03c15d2) on Aug 25, 2020 will the DSS enable automatic garbage collection of records by tracking which DSS instance is responsible for garbage collection of the record. Expired records added with a DSS deployment running code earlier than this must be manually removed. - -The Garbage collector job runs every 30 minute to delete records in RID tables that records' endtime is 30 minutes less than current time. If the event takes a long time and takes longer than 30 minutes (previous job is still running), the job will skip a run until the previous job completes. diff --git a/docs/infrastructure/index.md b/docs/infrastructure/index.md deleted file mode 100644 index d639d55f0..000000000 --- a/docs/infrastructure/index.md +++ /dev/null @@ -1,119 +0,0 @@ -# Introduction - -This section describes how to deploy a DSS instance on Kubernetes. - -## Deployment Options - -The DSS can be deployed on various platforms. Choose the method that best suits your needs: - -| Platform | Tools | Description | -| :--- | :--- | :--- | -| **Amazon Web Services** | Terraform | [Deploy on AWS using Terraform](aws.md) to provision EKS and required resources. | -| **Google Cloud Platform** | Terraform | [Deploy on GCP using Terraform](google.md) to provision GKE and required resources. | -| **Google Cloud Platform** | Manual | [Deploy on GCP manually](google-manual.md) without Terraform. | -| **Locally** | Minikube | [Deploy locally using Minikube](minikube.md) for development and testing. | - - -## Glossary - -- DSS Region - A region in which a single, unified airspace representation is - presented by one or more interoperable DSS instances, each instance typically - operated by a separate organization. A specific environment (for example, - "production" or "staging") in a particular DSS Region is called a "pool". -- DSS instance - a single logical replica in a DSS pool. - - -## Prerequisites - -Download & install the following tools to your workstation: - -- If deploying on Google Cloud, - [install Google Cloud SDK](https://cloud.google.com/sdk/install) - - Confirm successful installation with `gcloud version` - - Run `gcloud init` to set up a connection to your account. - - `kubectl` can be installed from `gcloud` instead of via the method below. -- [Install kubectl](https://kubernetes.io/docs/tasks/tools/install-kubectl/) to - interact with kubernetes - - Confirm successful installation with `kubectl version --client` (should - succeed from any working directory). - - Note that kubectl can alternatively be installed via the Google Cloud SDK - `gcloud` shell if using Google Cloud. -- [Install tanka](https://tanka.dev/install) - - On Linux, after downloading the binary per instructions, run - `sudo chmod +x /usr/local/bin/tk` - - Confirm successful installation with `tk --version` -- [Install Docker](https://docs.docker.com/get-docker/). - - Confirm successful installation with `docker --version` -- If using CockroachDB as the datastore, - [install CockroachDB](https://www.cockroachlabs.com/docs/v24.1/install-cockroachdb-linux) to - generate CockroachDB certificates. - - These instructions assume CockroachDB Core. - - You may need to run `sudo chmod +x /usr/local/bin/cockroach` after - completing the installation instructions. - - Confirm successful installation with `cockroach version` -- If developing the DSS codebase, - [install Golang](https://golang.org/doc/install) - - Confirm successful installation with `go version` -- Optionally install [Jsonnet](https://github.com/google/jsonnet) if editing - the jsonnet templates. - -## Docker images - -The application logic of the DSS is located in core-service which is provided in -a Docker image which is built locally and then pushed to a Docker registry of -your choice. All major cloud providers have a docker registry service, or you -can set up your own. - -To use the prebuilt InterUSS Docker images (without building them yourself), use -`docker.io/interuss/dss` for `VAR_DOCKER_IMAGE_NAME`. - -To build these images (and, optionally, push them to a docker registry): - -1. Set the environment variable `DOCKER_URL` to your docker registry url -endpoint. - - - For Google Cloud, `DOCKER_URL` should be set similarly to as described - [here](https://cloud.google.com/container-registry/docs/pushing-and-pulling#tag_the_local_image_with_the_registry_name), - like `gcr.io/your-project-id` (do not include the image name; - it will be appended by the build script) - - - For Amazon Web Services, `DOCKER_URL` should be set similarly to as described - [here](https://docs.aws.amazon.com/AmazonECR/latest/userguide/docker-push-ecr-image.html), - like `${aws_account_id}.dkr.ecr.${region}.amazonaws.com/` (do not include the image name; - it will be appended by the build script) - -1. Ensure you are logged into your docker registry service. - - - For Google Cloud, - [these](https://cloud.google.com/container-registry/docs/advanced-authentication#gcloud-helper) - are the recommended instructions (`gcloud auth configure-docker`). - Ensure that - [appropriate permissions are enabled](https://cloud.google.com/container-registry/docs/access-control). - - - For Amazon Web Services, create a private repository by following the instructions - [here](https://docs.aws.amazon.com/AmazonECR/latest/userguide/repository-create.html), then login - as described [here](https://docs.aws.amazon.com/AmazonECR/latest/userguide/docker-push-ecr-image.html). - -1. Use the [`build.sh` script](https://github.com/interuss/dss/blob/master/build/build.sh) in this directory to build and push - an image tagged with the current date and git commit hash. - -1. Note the VAR_* value printed at the end of the script. - -### Access to private repository - -See the description of `VAR_DOCKER_IMAGE_PULL_SECRET` to configure authentication [on the manual step by step guide](google-manual.md). - -### Verify signature of prebuilt InterUSS Docker images - -The prebuilt docker images are signed using [sigstore](https://www.sigstore.dev/). -The identity of the CI workflow, attested by GitHub, is used so sign the images. - -The signature may be verified by using [cosign](https://github.com/sigstore/cosign): -```shell -docker pull "docker.io/interuss/dss:latest" -cosign verify "docker.io/interuss/dss:latest" \ - --certificate-identity-regexp="https://github.com/interuss/dss/.github/workflows/dss-publish.yml@refs/*" \ - --certificate-oidc-issuer="https://token.actions.githubusercontent.com" -``` - -Adapt the version specified if required. diff --git a/docs/infrastructure/minikube.md b/docs/infrastructure/minikube.md deleted file mode 100644 index ea8f2467c..000000000 --- a/docs/infrastructure/minikube.md +++ /dev/null @@ -1,59 +0,0 @@ -# Deploy a DSS instance locally on Minikube - -This section provide instructions to prepare a local minikube cluster. - -Minikube is going to take care of most of the work by spawning a local kubernetes cluster. - -## Getting started - -### Prerequisites - -Download & install the following tools to your workstation: - -1. Install [minikube](https://minikube.sigs.k8s.io/docs/start/) (First step only). -2. Install tools from [Prerequisites](./google-manual.md#prerequisites) - -### Create a new minikube cluster - -1. Run `minikube start -p dss-local-cluster` to create a new cluster. -2. Run `minikube tunnel -p dss-local-cluster` and keep it running to expose LoadBalancer services. - -If needed, you can change the name of the cluster (`dss-local-cluster` in this documentation) as needed. You may also deploy multiple cluster at the same time, using different names. - -### Access to the cluster - -Minikube provide a UI, should you want to keep track of deployment and/or inspect the cluster. To start it, use the following command: - -1. `minikube dashboard -p dss-local-cluster` - -You can also use any other tool as needed. You can switch to the cluster's context by using the following command: - -1. `kubectl config use-context dss-local-cluster` - -### Upload or update local image - -Should you want to run the local docker image that you [built](./google-manual.md#prerequisites), run the following commands to upload / update your image - -1. `minikube image -p dss-local-cluster load interuss-local/dss` - -In the helm charts, use `docker.io/interuss-local/dss:latest` as image and be sure to set the `imagePullPolicy` to `Never`. - -## Deployment of the DSS services - -You can now deploy the DSS services using Helm or Tanka. See the repository `/deploy/services` for more information. - -=== "Helm" - Minikube specific settings: - - * Use the `global.cloudProvider` setting with the value `minikube` and deploy the charts on the `dss-local-cluster` kubernetes context. - -=== "Tanka" - An example configuration is provided in the repository: `/deploy/services/tanka/examples/minikube` - ---- - -To access the service, find the external IP using the `kubectl get services dss-dss-gateway` command. The port 80, without HTTPs is used. - -## Clean up - -To delete all resources, run `minikube delete -p dss-local-cluster`. Note that this operation can't be reverted and all data will be lost. diff --git a/docs/operations/.nav.yml b/docs/operations/.nav.yml index 3341a9ee7..46cb280e0 100644 --- a/docs/operations/.nav.yml +++ b/docs/operations/.nav.yml @@ -1,14 +1,12 @@ nav: - "Overview": index.md - - "Certificates management (Yugabyte)": certificates-management.md - - "Pooling (Yugabyte)": pooling.md - - "Pooling (CockroachDB)": pooling-crdb.md + - "Database cleanup": cleanup.md - "Monitoring": monitoring.md - "Health checks": healthchecks.md - - "Migrations": migrations.md - - "Upgrades": upgrades.md + - "Performance": performance.md - "Database migrations": database-migrations.md - - "Performances": performances.md - - "Cleanup": cleanup.md - - "Authentication": authentication.md + - "DSS upgrades": upgrades.md + - "CockroachDB version upgrades": crdb-upgrades.md + - "Kubernetes upgrades": kubernetes-upgrades.md + - "Pooling": pooling.md - "Troubleshooting": troubleshooting.md diff --git a/docs/operations/migrations.md b/docs/operations/crdb-upgrades.md similarity index 57% rename from docs/operations/migrations.md rename to docs/operations/crdb-upgrades.md index b2fbd084a..59b8b944a 100644 --- a/docs/operations/migrations.md +++ b/docs/operations/crdb-upgrades.md @@ -1,10 +1,8 @@ -# CockroachDB and Kubernetes version migration +# CockroachDB upgrades -This page provides information on how to upgrade your CockroachDB and Kubernetes cluster deployed using the +This page provides information on how to upgrade your CockroachDB version deployed using the tools from this repository. -## CockroachDB upgrades - CockroachDB must be upgraded on all DSS instances of the pool one after the other. The rollout of the upgrades on the whole CRDB cluster must be carefully performed in sequence to keep the majority of nodes healthy during that period and prevent downtime. @@ -26,13 +24,13 @@ be different by DSS instance. - We recommend to review carefully the instructions provided by CockroachDB and to rehearse all migrations on a test environment before applying them to production. -### Terraform deployment +## Terraform deployment -If a DSS instance has been deployed with terraform, first upgrade the cluster using [Helm](migrations.md#helm-deployment) -or [Tanka](migrations.md#tanka-deployment). Then, update the variable `crdb_image_tag` in your `terraform.tfvars` to +If a DSS instance has been deployed with terraform, first upgrade the cluster using [Helm](#helm-deployment) +or [Tanka](#tanka-deployment). Then, update the variable `crdb_image_tag` in your `terraform.tfvars` to align your configuration with the new state of the cluster. -### Helm deployment +## Helm deployment If you deployed the DSS using the Helm chart and the instructions provided in this repository, follow the instructions provided by CockroachDB `Cluster Upgrade with Helm` (See specific links below). Note that the CockroachDB documentation @@ -59,7 +57,7 @@ cockroachdb: New values can then be applied using `helm upgrade [RELEASE_NAME] [PATH_TO_DSS_HELM] -f [helm_values.yml]`. We recommend the second approach to keep your helm values in sync with the cluster state. -#### 21.2.7 to 24.1.3 +### 21.2.7 to 24.1.3 CockroachDB requires to upgrade one minor version at a time, therefore the following migrations have to be performed: @@ -69,15 +67,15 @@ CockroachDB requires to upgrade one minor version at a time, therefore the follo 1. 23.1 to 23.2: see [CockroachDB Cluster upgrade for Helm](https://www.cockroachlabs.com/docs/v23.2/upgrade-cockroachdb-kubernetes?filters=helm). 1. 23.2 to 24.1.3: see [CockroachDB Cluster upgrade for Helm](https://www.cockroachlabs.com/docs/v24.1/upgrade-cockroachdb-kubernetes?filters=helm). -### Tanka deployment +## Tanka deployment For deployments using Tanka configuration, since no instructions are provided for Tanka specifically, we recommend to follow the manual steps documented by CockroachDB: `Cluster Upgrade with Manual configs`. (See specific links below) To apply the changes to your cluster, follow the manual steps and reflect the new values in the *Leader* and *Followers* Tanka configurations, namely the new image version (see -[`VAR_CRDB_DOCKER_IMAGE_NAME`](../infrastructure/google-manual.md)) to ensure the new configuration is aligned with the cluster state. +`VAR_CRDB_DOCKER_IMAGE_NAME` in TFVARS.gen.md) to ensure the new configuration is aligned with the cluster state. -#### 21.2.7 to 24.1.3 +### 21.2.7 to 24.1.3 CockroachDB requires to upgrade one minor version at a time, therefore the following migrations have to be performed: @@ -86,69 +84,3 @@ CockroachDB requires to upgrade one minor version at a time, therefore the follo 1. 22.2 to 23.1: see [CockroachDB Cluster upgrade with Manual configs](https://www.cockroachlabs.com/docs/v23.1/upgrade-cockroachdb-kubernetes?filters=manual). 1. 23.1 to 23.2: see [CockroachDB Cluster upgrade with Manual configs](https://www.cockroachlabs.com/docs/v23.2/upgrade-cockroachdb-kubernetes?filters=manual). 1. 23.2 to 24.1.3: see [CockroachDB Cluster upgrade with Manual configs](https://www.cockroachlabs.com/docs/v24.1/upgrade-cockroachdb-kubernetes?filters=manual). - -## Kubernetes upgrades - -**Important notes:** - -- The migration plan below has been tested with the deployment of services using Helm and Tanka without Istio enabled. Note that this configuration flag has been decommissioned since [#995](https://github.com/interuss/dss/pull/995). -- Further work is required to test and evaluate the availability of the DSS during migrations. -- It is highly recommended to rehearse such operation on a test cluster before applying them to a production environment. - -### Google - Google Kubernetes Engine - -Migrations of GKE clusters are managed using terraform. - -#### 1.24 to 1.35 - -For each intermediate version up to the target version (eg. if you upgrade from 1.27 to 1.30, apply thoses -instructions for 1.28, 1.29, 1.30), do: - -Change your terraform.tfvars to use by adding or updating the kubernetes_version variable: -kubernetes_version = -Run terraform apply. This operation may take more than 30min. -Monitor the upgrade of the nodes in the Google Cloud console. - -1. Change your `terraform.tfvars` to use `` by adding or updating the `kubernetes_version` variable: - ```terraform - kubernetes_version = - ``` -1. Run `terraform apply`. This operation may take more than 30min. -1. Monitor the upgrade of the nodes in the Google Cloud console. - -### AWS - Elastic Kubernetes Service - -Currently, upgrades of EKS can't be achieved reliably with terraform directly. The recommended workaround is to -use the web console of AWS Elastic Kubernetes Service (EKS) to upgrade the cluster. -Before proceeding, always check on the cluster page the *Upgrade Insights* tab which provides a report of the -availability of Kubernetes resources in each version. The following sections omit this check if no resource is -expected to be reported in the context of a standard deployment performed with the tools in this repository. - -#### 1.25 to 1.35 - -1. Before migrating to 1.29, upgrade aws-load-balancer-controller helm chart on your cluster using `terraform apply`. Changes introduced by [PR #1167](https://github.com/interuss/dss/pull/1167). -You can verify if the operation has succeeded by running `helm list -n kube-system`. The APP VERSION shall be `2.12`. - -For each intermediate version up to the target version (eg. if you upgrade from 1.29 to 1.31, apply thoses -instructions for 1.29, 1.30, 1.31), do: -1. Upgrade the cluster (control plane) using the AWS console. It should take ~15 minutes. -1. Update the *Node Group* in the *Compute* tab with *Rolling Update* strategy to upgrade the nodes using the AWS console. - -To finalize the upgrade, change your `terraform.tfvars` to match the target version (ie 1.32) by adding or updating -the `kubernetes_version` variable: - ```terraform - kubernetes_version = 1.32 - ``` - -#### 1.24 to 1.25 - -1. Check for deprecated resources: - - Click on the Upgrade Insights tab to see deprecation warnings on the cluster page. - - Evaluate errors in Deprecated APIs removed in Kubernetes v1.25. Using `kubectl get podsecuritypolicies`, - check if there is only one *Pod Security Policy* named `eks.privileged`. If it is the case, - according to the [AWS documentation](https://docs.aws.amazon.com/eks/latest/userguide/pod-security-policy-removal-faq.html), you can proceed. -1. Upgrade the cluster using the AWS console. It should take ~15 minutes. -1. Change your `terraform.tfvars` to use `1.25` by adding or updating the `kubernetes_version` variable: - ```terraform - kubernetes_version = 1.25 - ``` diff --git a/docs/operations/index.md b/docs/operations/index.md index f66849934..d632e2639 100644 --- a/docs/operations/index.md +++ b/docs/operations/index.md @@ -1,45 +1,14 @@ # Operations -This section contains the instructions and related material used to operate a DSS. It is responsible to provide diagnostic capabilities and utilities to operate the DSS instance, such as certificates management. - -## Pooling procedure - -### Creating a new pool - -See [Creating a new pool](pooling.md#creating-a-new-pool) - -### Establishing a pool with first instance - -See [Establishing a pool with first instance](pooling.md#establishing-a-pool-with-first-instances) - -### Joining an existing pool with new instance - -See [Joining an existing pool with new instance](pooling.md#joining-an-existing-pool-with-new-instance) - -### Leaving a pool - -See [Leaving a pool](pooling.md#leaving-a-pool) - -## Monitoring - -See [Monitoring](monitoring.md) - -## Health checks - -See [Health checks](healthchecks.md) - -## Database migrations - -See [Database migrations](database-migrations.md) - -## Performances - -See [Performances](performances.md) - -## Authentication - -See [Authentication](authentication.md) - -## Troubleshooting - -See [Troubleshooting](troubleshooting.md) +This section contains the instructions and related material used to operate a deployed DSS instance. + +- [Database cleanup](cleanup.md) +- [Monitoring](monitoring.md) +- [Health checks](healthchecks.md) +- [Performance](performance.md) +- [Database migrations](database-migrations.md) +- [DSS upgrades](upgrades.md) +- [CockroachDB version upgrades](crdb-upgrades.md) +- [Kubernetes upgrades](kubernetes-upgrades.md) +- [Leaving a pool](pooling.md) +- [Troubleshooting](troubleshooting.md) diff --git a/docs/operations/kubernetes-upgrades.md b/docs/operations/kubernetes-upgrades.md new file mode 100644 index 000000000..b93d2928d --- /dev/null +++ b/docs/operations/kubernetes-upgrades.md @@ -0,0 +1,68 @@ +# Kubernetes upgrades + +This page provides information on how to upgrade your Kubernetes cluster deployed using the +tools from this repository. + +**Important notes:** + +- The migration plan below has been tested with the deployment of services using Helm and Tanka without Istio enabled. Note that this configuration flag has been decommissioned since [#995](https://github.com/interuss/dss/pull/995). +- Further work is required to test and evaluate the availability of the DSS during migrations. +- It is highly recommended to rehearse such operation on a test cluster before applying them to a production environment. + +## Google - Google Kubernetes Engine + +Migrations of GKE clusters are managed using terraform. + +### 1.24 to 1.35 + +For each intermediate version up to the target version (eg. if you upgrade from 1.27 to 1.30, apply thoses +instructions for 1.28, 1.29, 1.30), do: + +Change your terraform.tfvars to use by adding or updating the kubernetes_version variable: +kubernetes_version = +Run terraform apply. This operation may take more than 30min. +Monitor the upgrade of the nodes in the Google Cloud console. + +1. Change your `terraform.tfvars` to use `` by adding or updating the `kubernetes_version` variable: + ```terraform + kubernetes_version = + ``` +1. Run `terraform apply`. This operation may take more than 30min. +1. Monitor the upgrade of the nodes in the Google Cloud console. + +## AWS - Elastic Kubernetes Service + +Currently, upgrades of EKS can't be achieved reliably with terraform directly. The recommended workaround is to +use the web console of AWS Elastic Kubernetes Service (EKS) to upgrade the cluster. +Before proceeding, always check on the cluster page the *Upgrade Insights* tab which provides a report of the +availability of Kubernetes resources in each version. The following sections omit this check if no resource is +expected to be reported in the context of a standard deployment performed with the tools in this repository. + +### 1.25 to 1.35 + +1. Before migrating to 1.29, upgrade aws-load-balancer-controller helm chart on your cluster using `terraform apply`. Changes introduced by [PR #1167](https://github.com/interuss/dss/pull/1167). +You can verify if the operation has succeeded by running `helm list -n kube-system`. The APP VERSION shall be `2.12`. + +For each intermediate version up to the target version (eg. if you upgrade from 1.29 to 1.31, apply thoses +instructions for 1.29, 1.30, 1.31), do: +1. Upgrade the cluster (control plane) using the AWS console. It should take ~15 minutes. +1. Update the *Node Group* in the *Compute* tab with *Rolling Update* strategy to upgrade the nodes using the AWS console. + +To finalize the upgrade, change your `terraform.tfvars` to match the target version (ie 1.32) by adding or updating +the `kubernetes_version` variable: + ```terraform + kubernetes_version = 1.32 + ``` + +### 1.24 to 1.25 + +1. Check for deprecated resources: + - Click on the Upgrade Insights tab to see deprecation warnings on the cluster page. + - Evaluate errors in Deprecated APIs removed in Kubernetes v1.25. Using `kubectl get podsecuritypolicies`, + check if there is only one *Pod Security Policy* named `eks.privileged`. If it is the case, + according to the [AWS documentation](https://docs.aws.amazon.com/eks/latest/userguide/pod-security-policy-removal-faq.html), you can proceed. +1. Upgrade the cluster using the AWS console. It should take ~15 minutes. +1. Change your `terraform.tfvars` to use `1.25` by adding or updating the `kubernetes_version` variable: + ```terraform + kubernetes_version = 1.25 + ``` diff --git a/docs/operations/monitoring.md b/docs/operations/monitoring.md index 7c4aa508b..9becb309b 100644 --- a/docs/operations/monitoring.md +++ b/docs/operations/monitoring.md @@ -2,7 +2,7 @@ ## Prerequisites -Some of these [tools](../infrastructure/index.md#prerequisites) are required to interact with monitoring services. +Some of these [tools](../deployment/services/index.md#prerequisites) are required to interact with monitoring services. ## Grafana / Prometheus stack @@ -182,5 +182,5 @@ configured (use `admin` for both the username and password), Prometheus is available at http://localhost:9090, and Jaeger is available at http://localhost:16686. -See the [standalone local instance documentation](../../build/dev/standalone_instance.md#monitoring) +See the [standalone local instance documentation](https://github.com/interuss/dss/blob/master/build/dev/standalone_instance.md#monitoring) for more details. diff --git a/docs/operations/performances.md b/docs/operations/performance.md similarity index 99% rename from docs/operations/performances.md rename to docs/operations/performance.md index fe7135134..8c0a9ad59 100644 --- a/docs/operations/performances.md +++ b/docs/operations/performance.md @@ -1,4 +1,4 @@ -# Performances +# Performance ## Entries accumulation diff --git a/docs/operations/pooling.md b/docs/operations/pooling.md index dacd7cef5..ab375cf75 100644 --- a/docs/operations/pooling.md +++ b/docs/operations/pooling.md @@ -1,370 +1,7 @@ -# DSS Pooling (Yugabyte) +# Pooling operations -!!! note - This document is about pooling with **Yugabyte**. CockroachDB - documentation is [there](./pooling-crdb.md). +Pooling of a DSS instance (having it join an existing pool) is accomplished in the normal course of [deployment](../deployment/index.md). -## Introduction +An existing DSS instance cannot be moved to a different pool. Instead, the existing DSS instance must be decommissioned and a new DSS instance deployed to the different pool. -The DSS is designed to be deployed in a federated manner where multiple -organizations each host a DSS instance, and all of those instances interoperate. -Specifically, if a change is made on one DSS instance, that change may be read -from a different DSS instance. A set of interoperable DSS instances is called a -"pool", and the purpose of this document is to describe how to form and maintain -a DSS pool. - -It is expected that there will be exactly one production DSS pool for any given -DSS region, and that a DSS region will generally match aviation jurisdictional -boundaries (usually national boundaries). A given DSS region (e.g., -Switzerland) will likely have one pool for production operations, and an -additional pool for partner qualification and testing (per, e.g., -F3411-19 A2.6.2). - -### Terminology notes - -Yugabyte establishes a distributed data store called a "cluster". This cluster -stores the DSS Airspace Representation (DAR) in multiple SQL databases within -that cluster. This cluster is composed of many Yugabyte nodes, potentially -hosted by multiple organizations. - -Kubernetes manages a set of services in a "cluster". This is an entirely -different thing from the Yugabyte data store, and this type of cluster is what -the deployment instructions refer to. A Kubernetes cluster contains one or more -node pools: collections of machines available to run jobs. This node pool is an -entirely different thing from a DSS pool. - -## Objective - -A pool of InterUSS-compatible DSS instances is established when all of the -following requirements are met: - -1. Each Yugabyte node is addressable by every other Yugabyte node -1. Each Yugabyte node is discoverable -1. Each Yugabyte node accepts the certificates of every other node -1. The Yugabyte cluster is initialized - -The procedures in this document are intended to achieve all the objectives -listed above, but these procedures are not the only ways to achieve the -objectives. - -### "Each Yugabyte node is addressable by every other Yugabyte node" - -Every Yugabyte node must have its own externally-accessible hostname (e.g., -1.tserver.db.dss-prod.example.com), or its own hostname:port combination (e.g., -db.dss-prod.example.com:26258). - -There are two type of nodes in a Yugabyte cluster: Master and TServer. Both ones -must be accessible. The ports on which Yugabyte communicates must be open to -others participants: - -* Master: gRPC: **7100** -* TServer: gRPC: **9100** -* Master: Admin UI: 7000 -* TServer: Admin UI: 9000 -* TServer: ycql: 9042 -* TServer: ysql: 5433 -* TServer: metrics: 13000 -* TServer: metrics: 12000 - -The ports in bold are mandatory. The others ones are needed for management UI, -the UI won't work correctly if any of those port is not reachable by other -nodes, except on the master node. - -!!! info - The Helm charts and the Tanka files only expose mandatory ports as they are the - only ones secure. If usage of the UI is needed in a pool with multiple - participants, you must find a way to open those ports in a way secure enough - for your deployments. - Most of those non-mandatory ports do not offer authentication nor encryption - (or confidentiality). A secure method is required, such as an Istio mesh or - a local private network. - -This requirement may be verified by conducting a standard TLS diagnostic -(like [this one](https://www.wormly.com/test_ssl)) on the hostname:port -for each TServer node (e.g., 0.tserver.db.dss.example.com:7100). The "Trust" -characteristic will not pass because the certificate is issued by -a custom CA which is not a generally-trusted root CA, but we -explicitly enable trust by manually exchanging the trusted CA public keys -in ca.crt (see "Each Yugabyte node accepts the certificates of every other -node" below). However, all other checks should generally pass. - -!!! danger - It's recommended to restrict access to all ports and only allow IPs of - others participants. However, guides and deployment tooling haven't been adapted yet. - -### "Each Yugabyte node is discoverable" - -When a Yugabyte node is brought online, it must know how to connect to the -existing network of nodes. This is accomplished by providing an explicit list of -nodes to contact. Each node contacted will provide a list of nodes it is -connected to in the network ("gossip"), so not every node must be present in the -explicit list, but it's recommended to do so. The explicit list shall contain, -at a minimum, all known nodes when creating the DSS instance and shall be -updated regularly. Yugabyte nodes have some difficulites to locate primary nodes -if they don't have the full list of known nodes. - -### "Each Yugabyte node accepts the certificates of every other node" - -Yugabyte uses TLS to secure connections, and TLS includes a mechanism to -ensure the identity of the server being contacted. This mechanism requires a -trusted root Certificate Authority (CA) to sign a certificate containing the -public key of a particular server, so a client connecting to that server can -verify that the certificate (containing the public key) is endorsed by the root -CA as being genuine. Yugabyte certificates require a claim that standard web CAs -will not sign, so instead each USS acts as their own root CA. When USS 1 -is presented with certificates signed by USS 2's CA, USS 1 must know that it -can trust that certificate. We accomplish this by exchanging all USSs' CA -public keys out-of-band in ca.crt, and specifying that certificates signed by -any of the public keys in ca.crt should be accepted when considering the -validity of certificates presented to establish a TLS connection between nodes. - -The private CA key `dss-certs.py` generates is stored in the `ca` folder. The -private CA key is used to generate all node certificates and client -certificates. Once a pool is established, a USS avoids regenerating this CA -keypair, and use the existing ones by default. If a USS generates a new CA -keypair, the new public key must be added to the pool's combined ca.crt, and all -USSs in the pool must adopt the new combined ca.crt before any nodes using -certificates generated by the new CA private key will be accepted by the pool. - -### "The Yugabyte cluster is initialized" - -A Yugabyte cluster of databases is like the Ship of Theseus: it is composed of -many nodes which may all be replaced, one by one, so that a given Yugabyte -cluster eventually contains none of its original nodes. Unlike the Ship of -Theseus, however, a cluster is clearly identified by its cluster ID (e.g., -b2537de3-166f-42c4-aae1-742e094b8349) -- if the cluster ID is the same, it is -the same cluster (and vice versa). Once for the entire lifetime of the Yugabyte -cluster, it is created automatically during initialization of the first set of -Yugabyte nodes, if all initial nodes see each others in a uninitialized state. - -## Additional requirements - -These requirements must be met by every DSS instance joining an -InterUSS-compatible pool. The deployment instructions produce a system that -complies with all these requirements, so this section may be ignored if -following those instructions. - -- All Yugabyte nodes must be run in secure mode. - - use_node_to_node_encryption enabled - - use_client_to_server_encryption enabled - - node_to_node_encryption_use_client_certificates enabled - - allow_insecure_connections disabled -- All DSS instances in the same cluster must point their ntpd at the same NTP - Servers. - -## Creating a new pool -All DSS instances are equal peers, and any set of DSS instance can be chosen to -create the pool initially. After the pool is established, additional DSS -instance can join it. After that joining process is complete, it can be -repeated any number of times to add additional DSS instances, though 7 is the -maximum recommended number of DSS instances for performance reasons. The -following diagram illustrates the pooling process for the first two instances: - -![DSS pooling as a first, alone participant](../assets/generated/pool_new_1.png) - - -![DSS pooling with first 3 first participants](../assets/generated/pool_new_3.png) - -Adding participant is illustrated below. Some actions marked with `(once)` need -to be run only once by one participant otherwise all participants in the current -pool must ran then. - -![DSS pooling with new participant](../assets/generated/pool_add.png) - -### Establishing a pool with first instances -The USSs owning the first DSS instances should follow -[the deployment instructions](index.md). - -Each DSS instance must set `yugabyte_external_nodes` with the list of each -others DSS instance Yugabyte master nodes public endpoints, and CA certificates -must be exchanged. - -It's possible to have one DSS instance as starting point. In that case, -`yugabyte_external_nodes` will be empty and no CA exchange is needed. - -!!! info - Quick reminder for CA management: - - Each DSS instance should use `./dss-certs.sh init` To get the CA that should - be sent to others instances, use `./dss-certs.sh get-ca` To import the CA of - others DSS instance, use `./dss-certs.sh add-pool-ca` Finally, apply - certificates on the kubernetes cluster with `./dss-certs.sh apply` - -Ensure placement info is how you want it. See the section below for placement -requirements. - -Upon deployment completion, the following should be run against the DSS instance -to verify functionality: - - - The [prober test](https://github.com/interuss/monitoring/blob/main/monitoring/prober/README.md) - - The [USS qualifier](https://github.com/interuss/monitoring/tree/main/monitoring/uss_qualifier), - using the [DSS Probing](https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/configurations/dev/dss_probing.yaml) configuration - - -### Joining an existing pool with new instance -A USS wishing to join an existing pool (of perhaps just one instance following -the prior section) should follow [the deployment instructions](index.md). They -will be joining an existing cluster, and they will need to request all CAs that -the pool is currently using (any one member of the pool may provide it). The -joining USS will also need a list of Yugabyte node addresses. - -The joining USS must create his own CA with `./dss-certs.sh init` and retrieve -it with `./dss-certs.sh get-ca`. This certificate must be provided to each -existing DSS instance in the pool that will import it with `./dss-certs.sh -add-pool-ca` and `./dss-certs.sh apply`. - -One of existing DSS instance shall provide to the joining USS all existing -certificate, using `./dss-certs.sh get-pool-ca`. The joining USS can import them -with `./dss-certs.sh add-pool-ca` and finally apply certificates with -`./dss-certs.sh apply`. As an alternative, each DSS instance can provide its -individual CA. - -Participants shall ensure they work with a coherent set of certificate by -comparing the pool CA hash. It is displayed after adding certificates or using -the `./dss-certs.sh list-pool-ca`. - -When CAs have been exchanged and configured everywhere, the joining participant -can bring his system online (e.g. by applying helm charts onto his cluster). The -`yugabyte_external_nodes` setting shall be set **before** starting the Yugabyte -cluster. - -New nodes shall be allowed into the cluster. For each new Yugabyte master node, -the following command shall be run on one master node of one existing DSS -instance : - -!!! warning - The `master_addresses` in all commands below must include the Yugabyte master - leader. Either always run commands in the cluster with the leader, or list all - public addresses. - -1. Connection to a master node: - - `kubectl exec -it yb-master-0 -- sh` - -1. Addition of a new master node - - ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 change_master_config ADD_SERVER [PUBLIC HOSTNAME] 7100`` - -The last command can be repeated as needed, however a small delay is needed for -the cluster to settle when adding a new node. If you get `Leader is not ready -for Config Change, can try again`, just try again. - -You should have all masters listed in the web ui or using the -``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 list_all_masters`` -command. - -The tserver nodes will join automatically, using the list of provided master -nodes. They can be listed for confirmation in the web ui or using the -``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 list_all_tablet_servers`` -command. - -The pool should then be re-verified for functionality -by running the prober test on each DSS instance, and the -[interoperability test scenario](https://github.com/interuss/monitoring/blob/main/monitoring/uss_qualifier/scenarios/astm/netrid/v19/dss_interoperability.md) -on the full pool (including the newly-added instance). - -Finally, the joining USS should provide its Yugabyte node addresses to all other -participants in the pool, and each other participant should add those addresses -to the `yugabyte_external_nodes` list their Yugabyte nodes will attempt to -contact upon restart. - -Ensure placement info is how you want it. See the section below for placement -requirements. - -## Leaving a pool - -In an event that requires removing Yugabyte nodes we need to properly and -safely decommission to reduce risks of outages. - -It is never a good idea to take down more than half the number of nodes -available in your cluster as doing so would break quorum. If you need to take -down that many nodes please do it in smaller steps. - -Ensure placement info is how you want it after removal. Ensure you're not -requesting impossible placement by removing nodes, otherwise it won't be -possible to request node deletion. See the section below for placement -requirements. - -Note: If you are removing a specific node in a Statefulset, please know that -Kubernetes does not support removal of specific node; it automatically -re-creates the node if you delete it with `kubectl delete pod`. You will need -to scale down the Statefulset and that removes the last node first (ex: -`yb-tserver-n` where `n` is the `size of statefulset - 1`, `n` starts at 0) - -1. Check if all nodes are healthy in the web ui. - -1. Connect to a Yugabyte master and copy certs, like introduced in the previous - section. - -1. For each TServer node to be removed: - - 1. Blacklist one node in your cluster. - - ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 change_blacklist ADD [TSERVER_PUBLIC_HOSTNAME]`` - - 1. Wait for the node to be drained (no user tablet-peer or system-table-peer - in the gui). If node is not draining, you may have placement constraints - that prevent the removal of the node. - - 1. Stop one node in your cluster. - - 1. Wait until the node is marked as down and cluster will go into a - non-healthy state then wait for recovery. When everything is green again - proceed. Depending on settings, it may take time (15m) before the node is - marked as dead. - - 1. Remove the node: - - ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 remove_tablet_server [TSERVER_ID]`` - - If the command is giving you an error, data of the node may not have been - drain correctly dues to placement constraints. - - 1. Remove the node from the black list: - - ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 change_blacklist REMOVE [TSERVER_PUBLIC_HOSTNAME]`` - - 1. Fully remove the node in your cluster. - - E.g you may delete persistent volumes. - - -1. For each Master node to be removed: - - 1. Remove the master from the master list - - ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 change_master_config REMOVE_SERVER [PUBLIC HOSTNAME] 7100`` - - If the master node to be removed is the current leader, you may make it step - down with the following command: - - ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 master_leader_stepdown`` - -Finally, each pool participant should remove master addresses from the -`yugabyte_external_nodes` list their Yugabyte nodes will attempt to contact upon -restart and remove the CA of the participant. - -!!! note - Quick reminder for CA management: - - Remove the old CA, use `./dss-certs.sh remove-pool-ca ` - Finally, apply certificates on the kubernetes cluster with - `./dss-certs.sh apply` - -## Placement - -It's important to maintain a good placement strategy, ensuring data availability -in case of failures. - -We do recommend a minimum of 3 participants and one copy in each participants. - -You may use the `modify_placement_info` command to set placement settings. -Example: - -* ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 modify_placement_info dss.uss-1,dss.uss-2,dss.uss-3 3`` -* ``yb-admin -certs_dir_name /opt/certs/yugabyte/ -client_node_name=`hostname -f` -master_addresses yb-master-0.yb-masters.default.svc.cluster.local:7100,yb-master-1.yb-masters.default.svc.cluster.local:7100,yb-master-2.yb-masters.default.svc.cluster.local:7100 modify_placement_info dss.uss-1,dss.uss-2,dss.uss-3,dss.uss-4,dss.uss-5 5`` - -You may however use a different strategy depending on your availability needs, -e.g. you may want to avoid common datacenter between the same DSS instance. To -do so, define a strategy in your pool and edit placement information as needed. -More information is available in [Yugabyte documentation](https://docs.yugabyte.com/preview/admin/yb-admin/#modify-placement-info). +Depooling a DSS instance (removing it from the pool) is covered in [decommissioning](../decommissioning/index.md). diff --git a/docs/operations/upgrades.md b/docs/operations/upgrades.md index 01e783b1a..6e1d4ab08 100644 --- a/docs/operations/upgrades.md +++ b/docs/operations/upgrades.md @@ -47,7 +47,7 @@ of the versions listed above. ### Flags compatibility -Feature flags (particularly those detailed on [the performance page](./performances.md)) are +Feature flags (particularly those detailed on [the performance page](./performance.md)) are activated per node, with no global synchronization. Modifying these flags via a progressive rollout restart triggers a transitional state with the following impacts: diff --git a/mkdocs.yml b/mkdocs.yml index a044de754..5dd54c01a 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -1,6 +1,6 @@ # yaml-language-server: $schema=https://squidfunk.github.io/mkdocs-material/schema.json -site_name: DSS Deployment User Documentation +site_name: DSS User Documentation site_url: https://interuss.github.io/dss/ repo_url: https://github.com/interuss/dss repo_name: interuss/dss diff --git a/release/scripts/configure-clusters.sh b/release/scripts/configure-clusters.sh index 968be9f83..6be18b782 100755 --- a/release/scripts/configure-clusters.sh +++ b/release/scripts/configure-clusters.sh @@ -61,9 +61,9 @@ for ds in "${datastores[@]}"; do case "$ds" in ybdb) section "Yugabyte certs ($aws_name ↔ $goo_name)" - ( cd "$aws_ws" && ./dss-certs.sh init || true ) >"$aws_log" 2>&1 + ( cd "$aws_ws" && { ./dss-certs.sh init || true; } ) >"$aws_log" 2>&1 ok "init $aws_name" - ( cd "$goo_ws" && ./dss-certs.sh init || true ) >"$goo_log" 2>&1 + ( cd "$goo_ws" && { ./dss-certs.sh init || true; } ) >"$goo_log" 2>&1 ok "init $goo_name" ( cd "$goo_ws" && ./dss-certs.sh get-ca ) \