Skip to content

This is the multi-page printable view of this section. .

Return to the regular view of this page.

Internals

Discovery, logging, and Go module conventions for etcd contributors.

1 - Discovery service protocol

Discover other etcd members in a cluster bootstrap phase

Discovery service protocol helps new etcd member to discover all other members in cluster bootstrap phase using a shared discovery token and endpoint list.

Discovery service protocol is only used in cluster bootstrap phase, and cannot be used for runtime reconfiguration or cluster monitoring.

The protocol uses a new discovery token to bootstrap one unique etcd cluster. Remember that one discovery token can represent only one etcd cluster. As long as discovery protocol on this token starts, even if it fails halfway, it must not be used to bootstrap another etcd cluster.

The rest of this article will walk through the discovery process with examples that correspond to a self-hosted discovery cluster.

Note that this document is only for v3 discovery. Check previous document for more details on v2 discovery .

Protocol workflow

The idea of discovery protocol is to use an internal etcd cluster to coordinate bootstrap of a new cluster. First, all new members interact with discovery service and help to generate the expected member list. Then each new member bootstraps its server using this list, which performs the same functionality as -initial-cluster flag.

In the following example workflow, we will list each step of protocol using etcdctl command for ease of understanding, and we assume that http://example.com:2379 hosts an etcd cluster for discovery service.

By convention the etcd discovery protocol uses the key prefix /_etcd/registry.

Creating a new discovery token

Generate a unique token that will identify the new cluster. This will be used as a unique prefix in discovery keyspace in the following steps. An easy way to do this is to use uuidgen:

UUID=$(uuidgen)

Specifying the expected cluster size

The discovery token expects a cluster size that must be specified. The size is used by the discovery service to know when it has found all members that will initially form the cluster.

etcdctl --endpoints=http://example.com:2379 put /_etcd/registry/${UUID}/_config/size ${cluster_size}

Usually the cluster size is 3, 5 or 7. Check optimal cluster size for more details.

Bringing up etcd processes

Set the discovery token ${UUID} to --discovery-token flag, and set the endpoints of the etcd cluster backing the discovery service to --discovery-endpoints flag. This will enable v3 discovery to bootstrap the etcd cluster.

Every etcd process will follow the next few steps internally if --discovery-token and --discovery-endpoints flags are given.

If the discovery service enables client cert authentication, configure the following flags. They follow exactly the same usage as using etcdctl to communicate with an etcd cluster.

--discovery-insecure-transport
--discovery-insecure-skip-tls-verify
--discovery-cert
--discovery-key
--discovery-cacert

If the discovery service enables role based authentication, configure the following flags. They follow exactly the same usage as using etcdctl to communicate with an etcd cluster.

--discovery-user
--discovery-password

The default time or timeout values can also be changed using the following flags, which follow exactly the same usage as using etcdctl to communicate with an etcd cluster.

--discovery-dial-timeout
--discovery-request-timeout
--discovery-keepalive-time
--discovery-keepalive-timeout

Registering itself

The first thing that each etcd process does is to register itself into the given new cluster as a member. This is done by creating member ID as a key in the full registry key.

etcdctl --endpoints=http://example.com:2379 put /_etcd/registry/${UUID}/members/${member_id} ${member_name}=${member_peer_url_1}&${member_name}=${member_peer_url_2}

Checking the status

It checks the expected cluster size and registration status, and decides what the next action is.

etcdctl --endpoints=http://example.com:2379 get /_etcd/registry/${UUID}/_config/size
etcdctl --endpoints=http://example.com:2379 get /_etcd/registry/${UUID}/members

If registered members are still not enough, it will wait for other members to appear.

If the number of registered members is bigger than the expected size N, it treats the first N registered members as the member list for the cluster. If the member itself is in the member list, the discovery procedure succeeds, and it fetches all peers through the member list. If it is not in the member list, the discovery procedure finishes with the failure that the cluster has been full.

The member may check the cluster status even before registering itself. So it could fail quickly if the cluster has been full.

Waiting for all members

The wait process keeps watching the key prefix /_etcd/registry/${UUID}/members until finding all members.

etcdctl --endpoints=http://example.com:2379 watch /_etcd/registry/${UUID}/members --prefix

2 - Logging conventions

Logging level categories

etcd uses the zap library for logging application output categorized into levels. A log message’s level is determined according to these conventions:

  • DebugLevel logs are typically voluminous, and are usually disabled in production.

    • Examples:
      • Send a normal message to a remote peer
      • Write a log entry to disk
  • InfoLevel is the default logging priority.

    • Examples:
      • Startup configuration
      • Start to do snapshot
      • Add a new node into the cluster
      • Add a new user into auth subsystem
  • WarnLevel logs are more important than Info, but don’t need individual human review.

    • Examples:
      • Failure to send Raft message to a remote peer
      • Failure to receive heartbeat message within the configured election timeout
  • ErrorLevel logs are high-priority. If an application is running smoothly, it shouldn’t generate any error-level logs.

    • Examples:
      • Failure to allocate disk space for WAL
  • PanicLevel logs a message, then panics.

    • Examples:
      • Failure to encode Raft messages
  • FatalLevel logs a message, then calls os.Exit(1).

    • Examples:
      • Failure to save Raft snapshot

3 - Golang modules

Organization of the etcd project’s golang modules

The etcd project (since version 3.5) is organized into multiple golang modules hosted in a single repository .

modules graph

There are following modules:

  • go.etcd.io/etcd/api/v3 - contains API definitions (like protos & proto-generated libraries) that defines communication protocol between etcd clients and server.

  • go.etcd.io/etcd/pkg/v3 - collection of utility packages used by etcd without being specific to etcd itself. A package belongs here only if it could possibly be moved out into its own repository in the future. Please avoid adding here code that has a lot of dependencies on its own, as they automatically becoming dependencies of the client library (that we want to keep lightweight).

  • go.etcd.io/etcd/client/v3 - client library used to contact etcd over the network (grpc). Recommended for all new usage of etcd.

  • go.etcd.io/etcd/client/v2 - legacy client library used to contact etcd over HTTP protocol. Deprecated. All new usage should depend on /v3 library.

  • go.etcd.io/etcd/raft/v3 - implementation of distributed consensus protocol. Should have no etcd specific code.

  • go.etcd.io/etcd/server/v3 - etcd implementation. The code in this package is etcd internal and should not be consumed by external projects. The package layout and API can change within the minor versions.

  • go.etcd.io/etcd/etcdctl/v3 - a command line tool to access and manage etcd.

  • go.etcd.io/etcd/tests/v3 - a module that contains all integration tests of etcd. Notice: All unit-tests (fast and not requiring cross-module dependencies) should be kept in the local modules to the code under the test.

  • go.etcd.io/bbolt - implementation of persistent b-tree. Hosted in a separate repository: https://github.com/etcd-io/bbolt .

Operations

  1. All etcd modules should be released in the same versions, e.g. go.etcd.io/etcd/client/v3@v3.5.10 must depend on go.etcd.io/etcd/api/v3@v3.5.10.

    The consistent updating of versions can by performed using:

    % DRY_RUN=false TARGET_VERSION="v3.5.10" ./scripts/release_mod.sh update_versions
  2. The released modules should be tagged according to https://golang.org/ref/mod#vcs-version rules, i.e. each module should get its own tag. The tagging can be performed using:

    % DRY_RUN=false REMOTE_REPO="origin" ./scripts/release_mod.sh push_mod_tags
  3. All etcd modules should depend on the same versions of underlying dependencies. This can be verified using:

    % PASSES="dep" ./test.sh
  4. The go.mod files must not contain dependencies not being used and must conform to go mod tidy format. This is being verified by:

    % PASSES="mod_tidy" ./test.sh
  5. To trigger actions across all modules (e.g. auto-format all files), please use/expand the following script:

    % ./scripts/fix.sh

Future

As a North Star, we would like to evaluate etcd modules towards following model:

modules graph

This assumes:

  • Splitting etcdmigrate/etcdadm out of etcdctl binary. Thanks to this etcdctl would become clearly a command-line wrapper around network client API, while etcdmigrate/etcdadm would support direct physical operations on the etcd storage files.
  • Splitting etcd-proxy out of ./etcd binary, as it contains more experimental code so carries additional risk & dependencies.
  • Deprecation of support for v2 protocol.