exportETCDCTL_API=3ENDPOINTS=localhost:2379
etcdctl --endpoints=${ENDPOINTS} role add root
etcdctl --endpoints=${ENDPOINTS} role get root
etcdctl --endpoints=${ENDPOINTS} user add root
etcdctl --endpoints=${ENDPOINTS} user grant-role root root
etcdctl --endpoints=${ENDPOINTS} user get root
etcdctl --endpoints=${ENDPOINTS} role add role0
etcdctl --endpoints=${ENDPOINTS} role grant-permission role0 readwrite foo
etcdctl --endpoints=${ENDPOINTS} user add user0
etcdctl --endpoints=${ENDPOINTS} user grant-role user0 role0
etcdctl --endpoints=${ENDPOINTS} auth enable# now all client requests go through authetcdctl --endpoints=${ENDPOINTS} --user=user0:123 put foo bar
etcdctl --endpoints=${ENDPOINTS} get foo
# permission denied, user name is empty because the request does not issue an authentication requestetcdctl --endpoints=${ENDPOINTS} --user=user0:123 get foo
# user0 can read the key fooetcdctl --endpoints=${ENDPOINTS} --user=user0:123 get foo1
注意:
本文仅为示例,需补充并更新关于身份认证的更多信息。上述文本仅为代码示例。
1.2 - 基于角色的访问控制
一个基于角色的基本身份认证和访问控制指南
概述
身份认证功能自 etcd 2.1 版本起引入。etcd v3 API 对身份认证功能的 API 和用户界面进行了轻微调整,以更好地适配新的数据模型。本文旨在帮助用户在 etcd v3 中设置基本的身份认证和基于角色的访问控制。
etcd 不支持通过 --user username: 使用空密码进行身份认证。例如,使用空密码创建的用户,如 etcdctl user add anonymous:'',无法通过用户名/密码请求进行身份认证,类似 etcdctl --user anonymous: get foo 的请求将失败并返回 user name is empty。
用户的角色可使用以下方式授予或撤销:
$ etcdctl user grant-role myusername foo
$ etcdctl user revoke-role myusername bar
role 子命令用于 etcdctl,负责处理与特定角色访问控制相关的所有事项,这些权限已授予个别用户。
列出角色:
$ etcdctl role list
创建新角色,使用:
$ etcdctl role add myrolename
角色无密码;它仅用于定义一组新的访问权限。
角色被授予对单个键或键范围的访问权限。
范围可指定为区间 [起始键、结束键),其中起始键在字典序上应小于结束键。
访问权限可授予为读取、写入或两者兼有,例如以下示例所示:
# Give read access to a key /foo
$ etcdctl role grant-permission myrolename read /foo
# Give read access to keys with a prefix /foo/. The prefix is equal to the range [/foo/, /foo0)
$ etcdctl role grant-permission myrolename --prefix=true read /foo/
# Give write-only access to the key at /foo/bar
$ etcdctl role grant-permission myrolename write /foo/bar
# Give full access to keys in a range of [key1, key5)
$ etcdctl role grant-permission myrolename readwrite key1 key5
# Give full access to keys with a prefix /pub/
$ etcdctl role grant-permission myrolename --prefix=true readwrite /pub/
要查看已授予的权限,可随时查看角色:
$ etcdctl role get myrolename
权限撤销以相同逻辑方式进行:
$ etcdctl role revoke-permission myrolename /foo/bar
如移除角色本身:
$ etcdctl role delete myrolename
启用身份认证
启用身份认证的最小步骤如下。系统管理员可根据偏好,在启用身份认证之前或之后设置用户和角色。
确保已创建 root 用户:
$ etcdctl user add root
Password of root:
启用身份认证:
$ etcdctl auth enable
此后,etcd 已启用身份认证运行。如需出于任何原因禁用身份认证,请使用对应的反向命令:
$ etcdctl --user root:rootpw auth disable
身份认证的安全范围
当启用身份认证 etcdctl auth enable 时,可保护 V3 gRPC API 操作(get、put、delete、watch 等)。
--name 'default'
Human-readable name for this member.
--data-dir '${name}.etcd'
Path to the data directory.
--wal-dir ''
Path to the dedicated wal directory.
--snapshot-count '10000'
Number of committed transactions to trigger a snapshot to disk.
--heartbeat-interval '100'
Time (in milliseconds) of a heartbeat interval.
--election-timeout '1000'
Time (in milliseconds) for an election to timeout. See tuning documentation for details.
--initial-election-tick-advance 'true'
Whether to fast-forward initial election ticks on boot for faster election.
--listen-peer-urls 'http://localhost:2380'
List of URLs to listen on for peer traffic.
--listen-client-urls 'http://localhost:2379'
List of URLs to listen on for client grpc traffic and http as long as --listen-client-http-urls is not specified.
--listen-client-http-urls ''
List of URLs to listen on for http only client traffic. Enabling this flag removes http services from --listen-client-urls.
--max-snapshots '5'
Maximum number of snapshot files to retain (0 is unlimited).
--max-wals '5'
Maximum number of wal files to retain (0 is unlimited).
--memory-mlock
Enable to enforce etcd pages (in particular bbolt) to stay in RAM.
--quota-backend-bytes '0'
Raise alarms when backend size exceeds the given quota (0 defaults to low space quota).
--backend-bbolt-freelist-type 'map'
BackendFreelistType specifies the type of freelist that boltdb backend uses(array and map are supported types).
--backend-batch-interval ''
BackendBatchInterval is the maximum time before commit the backend transaction.
--backend-batch-limit '0'
BackendBatchLimit is the maximum operations before commit the backend transaction.
--max-txn-ops '128'
Maximum number of operations permitted in a transaction.
--max-request-bytes '1572864'
Maximum client request size in bytes the server will accept.
--grpc-keepalive-min-time '5s'
Minimum duration interval that a client should wait before pinging server.
--grpc-keepalive-interval '2h'
Frequency duration of server-to-client ping to check if a connection is alive (0 to disable).
--grpc-keepalive-timeout '20s'
Additional duration of wait before closing a non-responsive connection (0 to disable).
--socket-reuse-port 'false'
Enable to set socket option SO_REUSEPORT on listeners allowing rebinding of a port already in use.
--socket-reuse-address 'false'
Enable to set socket option SO_REUSEADDR on listeners allowing binding to an address in TIME_WAIT state.
集群管理
--initial-advertise-peer-urls 'http://localhost:2380'
List of this member's peer URLs to advertise to the rest of the cluster.
--initial-cluster 'default=http://localhost:2380'
Initial cluster configuration for bootstrapping.
--initial-cluster-state 'new'
Initial cluster state ('new' or 'existing').
--initial-cluster-token 'etcd-cluster'
Initial cluster token for the etcd cluster during bootstrap.
Specifying this can protect you from unintended cross-cluster interaction when running multiple clusters.
--advertise-client-urls 'http://localhost:2379'
List of this member's client URLs to advertise to the public.
The client URLs advertised should be accessible to machines that talk to etcd cluster. etcd client libraries parse these URLs to connect to the cluster.
--discovery ''
Discovery URL used to bootstrap the cluster.
--discovery-fallback 'proxy'
Expected behavior ('exit' or 'proxy') when discovery services fails.
"proxy" supports v2 API only.
--discovery-proxy ''
HTTP proxy to use for traffic to discovery service.
--discovery-srv ''
DNS srv domain used to bootstrap the cluster.
--discovery-srv-name ''
Suffix to the dns srv name queried when bootstrapping.
--strict-reconfig-check 'true'
Reject reconfiguration requests that would cause quorum loss.
--pre-vote 'true'
Enable the raft Pre-Vote algorithm to prevent disruption when a node that has been partitioned away rejoins the cluster.
--auto-compaction-retention '0'
Auto compaction retention length. 0 means disable auto compaction.
--auto-compaction-mode 'periodic'
Interpret 'auto-compaction-retention' one of: periodic|revision. 'periodic' for duration based retention, defaulting to hours if no time unit is provided (e.g. '5m'). 'revision' for revision number based retention.
--enable-v2 'false'
Accept etcd V2 client requests. Deprecated and to be decommissioned in v3.6.
--v2-deprecation 'not-yet'
Phase of v2store deprecation. Allows to opt-in for higher compatibility mode.
Supported values:
'not-yet' // Issues a warning if v2store have meaningful content (default in v3.5)
'write-only' // Custom v2 state is not allowed (default in v3.6 and v3.7)
'write-only-skip-check' // Custom v2 state is not supported and, if present, will be ignored (available in v3.5.32+, v3.6.13+, and v3.7.0+). Use this option at your own risk.
'write-only-drop-data' // Custom v2 state will get DELETED ! (planned default in v3.8)
'gone' // v2store is not maintained any longer.
安全
--cert-file ''
Path to the client server TLS cert file.
--key-file ''
Path to the client server TLS key file.
--client-cert-auth 'false'
Enable client cert authentication.
It's recommended to enable client cert authentication to prevent attacks from unauthenticated clients (e.g. CVE-2023-44487), especially when running etcd as a public service.
--client-crl-file ''
Path to the client certificate revocation list file.
--client-cert-allowed-hostname ''
Comma-separated list of SAN hostnames for client cert authentication.
--trusted-ca-file ''
Path to the client server TLS trusted CA cert file.
Note setting this parameter will also automatically enable client cert authentication no matter what value is set for `--client-cert-auth`.
--auto-tls 'false'
Client TLS using generated certificates.
--peer-cert-file ''
Path to the peer server TLS cert file.
--peer-key-file ''
Path to the peer server TLS key file.
--peer-client-cert-auth 'false'
Enable peer client cert authentication.
It's recommended to enable peer client cert authentication to prevent attacks from unauthenticated forged peers (e.g. CVE-2023-44487).
--peer-trusted-ca-file ''
Path to the peer server TLS trusted CA file.
--peer-cert-allowed-cn ''
Comma-separated list of allowed CNs for inter-peer TLS authentication.
--peer-cert-allowed-hostname ''
Comma-separated list of allowed SAN hostnames for inter-peer TLS authentication.
--peer-auto-tls 'false'
Peer TLS using self-generated certificates if --peer-key-file and --peer-cert-file are not provided.
--self-signed-cert-validity '1'
The validity period of the client and peer certificates that are automatically generated by etcd when you specify ClientAutoTLS and PeerAutoTLS, the unit is year, and the default is 1.
--peer-crl-file ''
Path to the peer certificate revocation list file.
--cipher-suites ''
Comma-separated list of supported TLS cipher suites between client/server and peers (empty will be auto-populated by Go).
--cors '*'
Comma-separated whitelist of origins for CORS, or cross-origin resource sharing, (empty or * means allow all).
--host-whitelist '*'
Acceptable hostnames from HTTP client requests, if server is not secure (empty or * means allow all).
--tls-min-version 'TLS1.2'
Minimum TLS version supported by etcd.
--tls-max-version ''
Maximum TLS version supported by etcd (empty will be auto-populated by Go).
认证
--auth-token 'simple'
Specify a v3 authentication token type and its options ('simple' or 'jwt').
--bcrypt-cost 10
Specify the cost / strength of the bcrypt algorithm for hashing auth passwords. Valid values are between 4 and 31.
--auth-token-ttl 300
Time (in seconds) of the auth-token-ttl.
性能分析和监控
--enable-pprof 'false'
Enable runtime profiling data via HTTP server. Address is at client URL + "/debug/pprof/"
--metrics 'basic'
Set level of detail for exported metrics, specify 'extensive' to include server side grpc histogram metrics.
--listen-metrics-urls ''
List of URLs to listen on for the metrics and health endpoints.
日志记录
--logger 'zap'
Currently only supports 'zap' for structured logging.
--log-outputs 'default'
Specify 'stdout' or 'stderr' to skip journald logging even when running under systemd, or list of comma separated output targets.
--log-level 'info'
Configures log level. Only supports debug, info, warn, error, panic, or fatal.
--log-format 'json'
Configures log format. Only supports json, console.
--enable-log-rotation 'false'
Enable log rotation of a single log-outputs file target.
--log-rotation-config-json '{"maxsize": 100, "maxage": 0, "maxbackups": 0, "localtime": false, "compress": false}'
Configures log rotation if enabled with a JSON logger config. MaxSize(MB), MaxAge(days,0=no limit), MaxBackups(0=no limit), LocalTime(use computers local time), Compress(gzip)".
--warning-unary-request-duration '300ms'
Set time duration after which a warning is logged if a unary request takes more than this duration.
--enable-distributed-tracing 'false'
Enable distributed tracing.
--distributed-tracing-address 'localhost:4317'
Distributed tracing collector address.
--distributed-tracing-service-name 'etcd'
Distributed tracing service name, must be the same across all etcd instances.
--distributed-tracing-instance-id ''
Distributed tracing instance ID, must be unique for each etcd instance.
--distributed-tracing-sampling-rate '0'
Number of samples to collect per million spans for distributed tracing.
v2 代理
警告
注意:标志位将在 v3.6 中被弃用。
--proxy 'off'
Proxy mode setting ('off', 'readonly' or 'on').
--proxy-failure-wait 5000
Time (in milliseconds) an endpoint will be held in a failed state.
--proxy-refresh-interval 30000
Time (in milliseconds) of the endpoints refresh interval.
--proxy-dial-timeout 1000
Time (in milliseconds) for a dial to timeout.
--proxy-write-timeout 5000
Time (in milliseconds) for a write to timeout.
--proxy-read-timeout 0
Time (in milliseconds) for a read to timeout.
功能
--corrupt-check-time '0s'
Duration of time between cluster corruption check passes.
--compact-hash-check-time '1m'
Duration of time between leader checks followers compaction hashes.
--compaction-batch-limit 1000
CompactionBatchLimit sets the maximum revisions deleted in each compaction batch.
--peer-skip-client-san-verification 'false'
Skip verification of SAN field in client certificate for peer connections.
--watch-progress-notify-interval '10m'
Duration of periodical watch progress notification.
--warning-apply-duration '100ms'
Warning is generated if requests take more than this duration.
--bootstrap-defrag-threshold-megabytes
Enable the defrag during etcd server bootstrap on condition that it will free at least the provided threshold of disk space. Needs to be set to non-zero value to take effect.
--max-learners '1'
Set the max number of learner members allowed in the cluster membership.
--compaction-sleep-interval
Sets the sleep interval between each compaction batch.
--downgrade-check-time
Duration of time between two downgrade status checks.
--snapshot-catchup-entries
Number of entries for a slow follower to catch up after compacting the raft storage entries.
功能门控
--feature-gates=AllAlpha=true|false
Enables or disables all alpha features. Default is false.
--feature-gates=AllBeta=true|false
Enables or disables all beta features. Default is false.
--feature-gates=CompactHashCheck=true
Enables leader to periodically check follower compaction hashes.
Replaces: --experimental-compact-hash-check-enabled
--feature-gates=InitialCorruptCheck=true
Enables corruption check before serving client/peer traffic.
Replaces: --experimental-initial-corrupt-check
--feature-gates=LeaseCheckpoint=true
ExperimentalEnableLeaseCheckpoint enables primary lessor to persist lease remainingTTL to prevent indefinite auto-renewal of long lived leases.
Replaces: --experimental-enable-lease-checkpoint
--feature-gates=LeaseCheckpointPersist=true
Enable persisting remainingTTL to prevent indefinite auto-renewal of long lived leases. Always enabled in v3.6. Should be used to ensure smooth upgrade from v3.5 clusters with this feature enabled.
Replaces: --experimental-enable-lease-checkpoint-persist
--feature-gates=SetMemberLocalAddr=true
Allows setting a member’s local address.
--feature-gates=StopGRPCServiceOnDefrag=true
Enable etcd gRPC service to stop serving client requests on defragmentation.
Replaces: --experimental-stop-grpc-service-on-defrag
--feature-gates=TxnModeWriteWithSharedBuffer=true
Enable the write transaction to use a shared buffer in its readonly check operations.
Replaces: --experimental-txn-mode-write-with-shared-buffer
不安全功能
警告
警告:使用不安全功能可能会破坏共识协议所提供的保证!
--force-new-cluster 'false'
Force to create a new one-member cluster.
--unsafe-no-fsync 'false'
Disables fsync, unsafe, will cause data loss.
自 v3.2.0
起,服务器拒绝包含错误 IP 地址的对等成员证书 SAN
。例如,若对等成员证书的 Subject Alternative Name (SAN) 字段中包含任何 IP 地址,服务器仅在远程 IP 地址与其中任一 IP 地址匹配时才认证该对等成员。此举旨在防止未经授权的端点加入集群。例如,对等成员 B 的 CSR(含 cfssl)为:
当对等成员 B 的实际 IP 地址为 10.138.0.2 时,而非 10.138.0.27。当对等成员 B 尝试加入集群时,对等成员 A 将以错误 x509: certificate is valid for 10.138.0.27, not 10.138.0.2 拒绝 B 的加入请求,因为 B 的远程 IP 地址与 Subject Alternative Name (SAN) 字段中的地址不匹配。
自 v3.2.0
起,服务器在检查 SAN
时会解析 TLS DNSNames。例如,若对等成员证书的 Subject Alternative Name(SAN)字段中仅包含 DNS 名称(无 IP 地址),则服务器仅在对这些 DNS 名称执行正向查找(dig b.com)并确认其解析出的 IP 地址与远程 IP 地址匹配时,才完成对等成员的身份认证。例如,对等成员 B 的 CSR(含 cfssl)为:
{"CN":"etcd peer","hosts":["b.com"],
当对等成员 B 的远程 IP 地址为 10.138.0.2 时。当对等成员 B 尝试加入集群时,对等成员 A 会查找入站主机 b.com 以获取 IP 地址列表(例如 dig b.com)。如果该列表不包含 IP 10.138.0.2,则拒绝 B 的加入请求,并返回错误 tls: 10.138.0.2 does not match any of DNSNames ["b.com"]。
自 v3.2.2
起,服务器在 IP 地址匹配时接受连接,不再检查 DNS 条目
。例如,若对等成员证书中的 Subject Alternative Name (SAN) 字段包含 IP 地址和 DNS 名称,且远程 IP 地址与其中任一 IP 地址匹配,服务器将直接接受连接,不再进一步验证 DNS 名称。例如,对等成员 B 的 CSR(含 cfssl)为:
当对等成员 B 的远程 IP 地址为 10.138.0.2 且 invalid.domain 为无效主机时,对等成员 B 尝试加入集群,对等成员 A 可成功对 B 进行身份认证,因为主题备用名称(SAN)字段包含有效的匹配 IP 地址。详情请参见 issue#8206
。
自 v3.2.5
起,服务器支持对通配符 DNS 的反向查找 SAN
。例如,若对等成员证书中的 Subject Alternative Name(SAN)字段仅包含 DNS 名称(无 IP 地址),服务器首先对远程 IP 地址执行反向查找,以获取映射到该地址的一组名称(例如 nslookup IPADDR)。若这些名称中存在与对等成员证书中 DNS 名称匹配的名称(通过精确匹配或通配符匹配),则接受连接。若无匹配项,服务器将对证书中的每个 DNS 条目执行正向查找(例如,当条目为 *.example.default.svc 时,查找 example.default.svc),仅当主机解析出的地址中包含与对等成员远程 IP 地址匹配的 IP 地址时,才接受连接。例如,对等成员 B 的 CSR(含 cfssl)为:
当对等成员 B 的远程 IP 地址为 10.138.0.2 时。对等成员 B 尝试加入集群,对等成员 A 会反向查找 IP 10.138.0.2 以获取主机名列表,并将主机名与对等成员 B 证书中 Subject Alternative Name (SAN) 字段的 DNS 名称进行精确匹配或通配符匹配。若反向或正向查找均失败,将返回错误 "tls: "10.138.0.2" does not match any of DNSNames ["*.example.default.svc","*.example.default.svc.cluster.local"]。详情请参见 issue#8268
。
v3.3.0
引入 etcd --peer-cert-allowed-cn
标志,以支持对等成员间连接的 基于 CN(通用名称)的身份认证
。Kubernetes TLS 引导机制涉及为 etcd 成员及其他系统组件(例如 API 服务器、kubelet 等)生成动态证书。为每个组件维护不同的 CA 可提供对 etcd 集群更严格的访问控制,但通常较为繁琐。当指定 –peer-cert-allowed-cn 标志时,节点仅能以匹配的通用名称加入集群,即使使用共享 CA 亦然。匹配方式为与证书的通用名称(CN)字段进行精确字符串比较——不支持通配符或前缀匹配。对于基于主机名的过滤,使用 –peer-cert-allowed-hostname 或 –client-cert-allowed-hostname 时,匹配采用 Go 的 x509.Certificate.VerifyHostname() 函数,支持精确主机名及通配符条目(例如 *.example.com)。例如,三节点集群中每个成员使用 CSRs(通过 cfssl)配置如下:
若提供了 --peer-cert-allowed-cn etcd.local,则仅对等成员中 Common Name 匹配的才会被认证。若证书签名请求(CSR)中的 CN 不同,或 --peer-cert-allowed-cn 不同,则节点将被拒绝:
$ etcd --peer-cert-allowed-cn m1.etcd.local
I | embed: rejected connection from "127.0.0.1:48044"(error "CommonName authentication failed", ServerName "m1.etcd.local")I | embed: rejected connection from "127.0.0.1:55702"(error "remote error: tls: bad certificate", ServerName "m3.etcd.local")
每个进程应以以下方式启动:
etcd --peer-cert-allowed-cn etcd.local
I | pkg/netutil: resolving m3.etcd.local:32380 to 127.0.0.1:32380
I | pkg/netutil: resolving m2.etcd.local:22380 to 127.0.0.1:22380
I | pkg/netutil: resolving m1.etcd.local:2380 to 127.0.0.1:2380
I | etcdserver: published {Name:m3 ClientURLs:[https://m3.etcd.local:32379]} to cluster 9db03f09b20de32b
I | embed: ready to serve client requests
I | etcdserver: published {Name:m1 ClientURLs:[https://m1.etcd.local:2379]} to cluster 9db03f09b20de32b
I | embed: ready to serve client requests
I | etcdserver: published {Name:m2 ClientURLs:[https://m2.etcd.local:22379]} to cluster 9db03f09b20de32b
I | embed: ready to serve client requests
I | embed: serving client requests on 127.0.0.1:32379
I | embed: serving client requests on 127.0.0.1:22379
I | embed: serving client requests on 127.0.0.1:2379
$ dig +noall +answer SRV _etcd-server._tcp.example.com
_etcd-server._tcp.example.com. 300 IN SRV 0 0 2380 infra0.example.com.
_etcd-server._tcp.example.com. 300 IN SRV 0 0 2380 infra1.example.com.
_etcd-server._tcp.example.com. 300 IN SRV 0 0 2380 infra2.example.com.
$ dig +noall +answer SRV _etcd-client._tcp.example.com
_etcd-client._tcp.example.com. 300 IN SRV 0 0 2379 infra0.example.com.
_etcd-client._tcp.example.com. 300 IN SRV 0 0 2379 infra1.example.com.
_etcd-client._tcp.example.com. 300 IN SRV 0 0 2379 infra2.example.com.
$ dig +noall +answer infra0.example.com infra1.example.com infra2.example.com
infra0.example.com. 300 IN A 10.0.1.10
infra1.example.com. 300 IN A 10.0.1.11
infra2.example.com. 300 IN A 10.0.1.12
使用 DNS 引导 etcd 集群
etcd 集群成员可通告域名或 IP 地址,引导过程将解析 DNS A 记录。
自 3.2 版本起(3.1 版本会打印警告)--listen-peer-urls 和 --listen-client-urls 将拒绝为网络接口绑定使用域名。
# file: etcd.yaml---apiVersion:v1kind:Servicemetadata:name:etcdnamespace:defaultspec:type:ClusterIPclusterIP:Noneselector:app:etcd#### Ideally we would use SRV records to do peer discovery for initialization.## Unfortunately discovery will not work without logic to wait for these to## populate in the container. This problem is relatively easy to overcome by## making changes to prevent the etcd process from starting until the records## have populated. The documentation on statefulsets briefly talk about it.## https://kubernetes.io/docs/concepts/workloads/controllers/statefulset/#stable-network-idpublishNotReadyAddresses:true#### The naming scheme of the client and server ports match the scheme that etcd## uses when doing discovery with SRV records.ports:- name:etcd-clientport:2379- name:etcd-serverport:2380- name:etcd-metricsport:8080---apiVersion:apps/v1kind:StatefulSetmetadata:namespace:defaultname:etcdspec:#### The service name is being set to leverage the service headlessly.## https://kubernetes.io/docs/concepts/services-networking/service/#headless-servicesserviceName:etcd#### If you are increasing the replica count of an existing cluster, you should## also update the --initial-cluster-state flag as noted further down in the## container configuration.replicas:3#### For initialization, the etcd pods must be available to eachother before## they are "ready" for traffic. The "Parallel" policy makes this possible.podManagementPolicy:Parallel#### To ensure availability of the etcd cluster, the rolling update strategy## is used. For availability, there must be at least 51% of the etcd nodes## online at any given time.updateStrategy:type:RollingUpdate#### This is label query over pods that should match the replica count.## It must match the pod template's labels. For more information, see the## following documentation:## https://kubernetes.io/docs/concepts/overview/working-with-objects/labels/#label-selectorsselector:matchLabels:app:etcd#### Pod configuration template.template:metadata:#### The labeling here is tied to the "matchLabels" of this StatefulSet and## "affinity" configuration of the pod that will be created.#### This example's labeling scheme is fine for one etcd cluster per## namespace, but should you desire multiple clusters per namespace, you## will need to update the labeling schema to be unique per etcd cluster.labels:app:etcdannotations:#### This gets referenced in the etcd container's configuration as part of## the DNS name. It must match the service name created for the etcd## cluster. The choice to place it in an annotation instead of the env## settings is because there should only be 1 service per etcd cluster.serviceName:etcdspec:#### Configuring the node affinity is necessary to prevent etcd servers from## ending up on the same hardware together.#### See the scheduling documentation for more information about this:## https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#node-affinityaffinity:## The podAntiAffinity is a set of rules for scheduling that describe## when NOT to place a pod from this StatefulSet on a node.podAntiAffinity:#### When preparing to place the pod on a node, the scheduler will check## for other pods matching the rules described by the labelSelector## separated by the chosen topology key.requiredDuringSchedulingIgnoredDuringExecution:## This label selector is looking for app=etcd- labelSelector:matchExpressions:- key:appoperator:Invalues:- etcd## This topology key denotes a common label used on nodes in the## cluster. The podAntiAffinity configuration essentially states## that if another pod has a label of app=etcd on the node, the## scheduler should not place another pod on the node.## https://kubernetes.io/docs/reference/labels-annotations-taints/#kubernetesiohostnametopologyKey:"kubernetes.io/hostname"#### Containers in the podcontainers:## This example only has this etcd container.- name:etcdimage:quay.io/coreos/etcd:v3.7.0imagePullPolicy:IfNotPresentports:- name:etcd-clientcontainerPort:2379- name:etcd-servercontainerPort:2380- name:etcd-metricscontainerPort:8080#### These probes will fail over TLS for self-signed certificates, so etcd## is configured to deliver metrics over port 8080 further down.#### As mentioned in the "Monitoring etcd" page, /readyz and /livez were## added in v3.5.12. Prior to this, monitoring required extra tooling## inside the container to make these probes work.#### The values in this readiness probe should be further validated, it## is only an example configuration.readinessProbe:httpGet:path:/readyzport:8080initialDelaySeconds:10periodSeconds:5timeoutSeconds:5successThreshold:1failureThreshold:30## The values in this liveness probe should be further validated, it## is only an example configuration.livenessProbe:httpGet:path:/livezport:8080initialDelaySeconds:15periodSeconds:10timeoutSeconds:5failureThreshold:3env:#### Environment variables defined here can be used by other parts of the## container configuration. They are interpreted by Kubernetes, instead## of in the container environment.#### These env vars pass along information about the pod.- name:K8S_NAMESPACEvalueFrom:fieldRef:fieldPath:metadata.namespace- name:HOSTNAMEvalueFrom:fieldRef:fieldPath:metadata.name- name:SERVICE_NAMEvalueFrom:fieldRef:fieldPath:metadata.annotations['serviceName']#### Configuring etcdctl inside the container to connect to the etcd node## in the container reduces confusion when debugging.- name:ETCDCTL_ENDPOINTSvalue:$(HOSTNAME).$(SERVICE_NAME):2379#### TLS client configuration for etcdctl in the container.## These files paths are part of the "etcd-client-certs" volume mount.# - name: ETCDCTL_KEY# value: /etc/etcd/certs/client/tls.key# - name: ETCDCTL_CERT# value: /etc/etcd/certs/client/tls.crt# - name: ETCDCTL_CACERT# value: /etc/etcd/certs/client/ca.crt#### Use this URI_SCHEME value for non-TLS clusters.- name:URI_SCHEMEvalue:"http"## TLS: Use this URI_SCHEME for TLS clusters.# - name: URI_SCHEME# value: "https"#### If you're using a different container, the executable may be in a## different location. This example uses the full path to help remove## ambiguity to you, the reader.## Often you can just use "etcd" instead of "/usr/local/bin/etcd" and it## will work because the $PATH includes a directory containing "etcd".command:- /usr/local/bin/etcd#### Arguments used with the etcd command inside the container.args:#### Configure the name of the etcd server.- --name=$(HOSTNAME)#### Configure etcd to use the persistent storage configured below.- --data-dir=/data#### In this example we're consolidating the WAL into sharing space with## the data directory. This is not ideal in production environments and## should be placed in it's own volume.- --wal-dir=/data/wal#### URL configurations are parameterized here and you shouldn't need to## do anything with these.- --listen-peer-urls=$(URI_SCHEME)://0.0.0.0:2380- --listen-client-urls=$(URI_SCHEME)://0.0.0.0:2379- --advertise-client-urls=$(URI_SCHEME)://$(HOSTNAME).$(SERVICE_NAME):2379#### This must be set to "new" for initial cluster bootstrapping. To scale## the cluster up, this should be changed to "existing" when the replica## count is increased. If set incorrectly, etcd makes an attempt to## start but fail safely.- --initial-cluster-state=new#### Token used for cluster initialization. The recommendation for this is## to use a unique token for every cluster. This example parameterized## to be unique to the namespace, but if you are deploying multiple etcd## clusters in the same namespace, you should do something extra to## ensure uniqueness amongst clusters.- --initial-cluster-token=etcd-$(K8S_NAMESPACE)#### The initial cluster flag needs to be updated to match the number of## replicas configured. When combined, these are a little hard to read.## Here is what a single parameterized peer looks like:## etcd-0=$(URI_SCHEME)://etcd-0.$(SERVICE_NAME):2380- --initial-cluster=etcd-0=$(URI_SCHEME)://etcd-0.$(SERVICE_NAME):2380,etcd-1=$(URI_SCHEME)://etcd-1.$(SERVICE_NAME):2380,etcd-2=$(URI_SCHEME)://etcd-2.$(SERVICE_NAME):2380#### The peer urls flag should be fine as-is.- --initial-advertise-peer-urls=$(URI_SCHEME)://$(HOSTNAME).$(SERVICE_NAME):2380#### This avoids probe failure if you opt to configure TLS.- --listen-metrics-urls=http://0.0.0.0:8080#### These are some configurations you may want to consider enabling, but## should look into further to identify what settings are best for you.# - --auto-compaction-mode=periodic# - --auto-compaction-retention=10m#### TLS client configuration for etcd, reusing the etcdctl env vars.# - --client-cert-auth# - --trusted-ca-file=$(ETCDCTL_CACERT)# - --cert-file=$(ETCDCTL_CERT)# - --key-file=$(ETCDCTL_KEY)#### TLS server configuration for etcdctl in the container.## These files paths are part of the "etcd-server-certs" volume mount.# - --peer-client-cert-auth# - --peer-trusted-ca-file=/etc/etcd/certs/server/ca.crt# - --peer-cert-file=/etc/etcd/certs/server/tls.crt# - --peer-key-file=/etc/etcd/certs/server/tls.key#### This is the mount configuration.volumeMounts:- name:etcd-datamountPath:/data#### TLS client configuration for etcdctl# - name: etcd-client-tls# mountPath: "/etc/etcd/certs/client"# readOnly: true#### TLS server configuration# - name: etcd-server-tls# mountPath: "/etc/etcd/certs/server"# readOnly: truevolumes:#### TLS client configuration# - name: etcd-client-tls# secret:# secretName: etcd-client-tls# optional: false#### TLS server configuration# - name: etcd-server-tls# secret:# secretName: etcd-server-tls# optional: false#### This StatefulSet will uses the volumeClaimTemplate field to create a PVC in## the cluster for each replica. These PVCs can not be easily resized later.volumeClaimTemplates:- metadata:name:etcd-dataspec:accessModes:["ReadWriteOnce"]#### In some clusters, it is necessary to explicitly set the storage class.## This example will end up using the default storage class.# storageClassName: ""resources:requests:storage:1Gi
高级集群管理系统(如 Kubernetes)原生支持服务发现。应用程序可通过系统管理的 DNS 名称或虚拟 IP 地址访问 etcd 集群。例如,kube-proxy 相当于 etcd 网关。
启动 etcd 网关
考虑一个具有以下静态端点的 etcd 集群:
名称
地址
主机名
端口
infra0
10.0.1.10
infra0.example.com
2379
infra1
10.0.1.11
infra1.example.com
2379
infra2
10.0.1.12
infra2.example.com
2379
使用以下命令启动 etcd 网关,以通过静态端点进行访问:
$ etcd gateway start --endpoints=infra0.example.com:2379,infra1.example.com:2379,infra2.example.com:2379
2016-08-16 11:21:18.867350 I | tcpproxy: ready to proxy client requests to [...]
或者,若使用 DNS 进行服务发现,请考虑使用 DNS SRV 记录:
$ dig +noall +answer SRV _etcd-client._tcp.example.com
_etcd-client._tcp.example.com. 300 IN SRV 002379 infra0.example.com.
_etcd-client._tcp.example.com. 300 IN SRV 002379 infra1.example.com.
_etcd-client._tcp.example.com. 300 IN SRV 002379 infra2.example.com.
$ dig +noall +answer infra0.example.com infra1.example.com infra2.example.com
infra0.example.com. 300 IN A 10.0.1.10
infra1.example.com. 300 IN A 10.0.1.11
infra2.example.com. 300 IN A 10.0.1.12
使用以下命令启动 etcd 网关,从 DNS SRV 条目中获取端点:
$ etcd gateway start --discovery-srv=example.com
2016-08-16 11:21:18.867350 I | tcpproxy: ready to proxy client requests to [...]
ETCDCTL_API=3 etcdctl --endpoints=http://localhost:23790 member list --write-out table
+----+---------+--------------------------------+------------+-----------------+
| ID | STATUS | NAME | PEER ADDRS | CLIENT ADDRS |+----+---------+--------------------------------+------------+-----------------+
|0| started | Gyu-Hos-MBP.sfo.coreos.systems || 127.0.0.1:23791 ||0| started | Gyu-Hos-MBP.sfo.coreos.systems || 127.0.0.1:23790 |+----+---------+--------------------------------+------------+-----------------+
ETCDCTL_API=3 etcdctl --endpoints=http://localhost:23792 member list --write-out table
+----+---------+--------------------------------+------------+-----------------+
| ID | STATUS | NAME | PEER ADDRS | CLIENT ADDRS |+----+---------+--------------------------------+------------+-----------------+
|0| started | Gyu-Hos-MBP.sfo.coreos.systems || 127.0.0.1:23792 |+----+---------+--------------------------------+------------+-----------------+
$ curl -L http://localhost:2379/metrics | grep -v debugging # ignore unstable debugging metrics# HELP etcd_disk_backend_commit_duration_seconds The latency distributions of commit called by backend.# TYPE etcd_disk_backend_commit_duration_seconds histogrametcd_disk_backend_commit_duration_seconds_bucket{le="0.002"}72756etcd_disk_backend_commit_duration_seconds_bucket{le="0.004"}401587etcd_disk_backend_commit_duration_seconds_bucket{le="0.008"}405979etcd_disk_backend_commit_duration_seconds_bucket{le="0.016"}406464...
$ etcdctl member add infra3 --peer-urls=http://10.0.1.13:2380
added member 9bf1b35fc7761a23 to cluster
ETCD_NAME="infra3"ETCD_INITIAL_CLUSTER="infra0=http://10.0.1.10:2380,infra1=http://10.0.1.11:2380,infra2=http://10.0.1.12:2380,infra3=http://10.0.1.13:2380"ETCD_INITIAL_CLUSTER_STATE=existing
若添加多个成员,最佳实践是逐个配置成员,并在添加更多新成员前验证每个成员是否已正确启动。若向单成员集群添加新成员,在新成员启动前,集群无法推进,因为达成共识需要多数成员(即至少两个成员)达成一致。此行为仅发生在 etcdctl member add 通知集群新成员存在,且新成员成功与现有成员建立连接之间的时段。
使用 etcdctl member add 并配合标志 --learner,可将新成员作为学习者成员添加至集群。
$ etcdctl member add infra3 --peer-urls=http://10.0.1.13:2380 --learner
Member 9bf1b35fc7761a23 added to cluster a7ef944b95711739
ETCD_NAME="infra3"ETCD_INITIAL_CLUSTER="infra0=http://10.0.1.10:2380,infra1=http://10.0.1.11:2380,infra2=http://10.0.1.12:2380,infra3=http://10.0.1.13:2380"ETCD_INITIAL_CLUSTER_STATE=existing
新添加的学习者成员启动新的 etcd 进程后,使用 etcdctl member promote 将该学习者成员提升为投票成员。
$ etcdctl member promote 9bf1b35fc7761a23
Member 9e29bbaa45d74461 promoted in cluster a7ef944b95711739