kubernetes

Author	SHA1	Message	Date
vinay kulkarni	01b96e7704	Rename ContainerStatus.ResourcesAllocated to ContainerStatus.AllocatedResources	2023-03-10 14:49:26 +00:00
Kubernetes Prow Robot	f734741cb8	Merge pull request #114373 from TommyStarK/unit-tests/kubelet-kuberuntime kubelet/kuberuntime: Improving test coverage	2023-03-10 04:34:58 -08:00
TommyStarK	7f21a9ce01	kubelet/kuberuntime: Improving test coverage Signed-off-by: TommyStarK <thomasmilox@gmail.com>	2023-03-10 11:06:54 +01:00
Kubernetes Prow Robot	3219564cf3	Merge pull request #116296 from SataQiu/clean-kubelet-20230306 Remove unused resize.go from pkg/kubelet/container	2023-03-09 22:43:48 -08:00
Kubernetes Prow Robot	e57d968323	Merge pull request #116015 from SataQiu/clean-kubelet-20230223 kubelet: remove the deprecated --master-service-namespace flag	2023-03-09 22:43:34 -08:00
Kubernetes Prow Robot	a408be817f	Merge pull request #115972 from jsafrane/add-orphan-pod-metrics Add metric for failed orphan pod cleanup	2023-03-09 22:43:26 -08:00
Kubernetes Prow Robot	33d8614c9c	Merge pull request #115929 from HirazawaUi/delete-kubelet-unused-function cleanup(kubelet): remove unused function	2023-03-09 22:43:12 -08:00
Kubernetes Prow Robot	0018c07050	Merge pull request #115898 from saschagrunert/seccomp-todo Default to sandbox `Seccomp` field instead of `SeccompProfilePath`	2023-03-09 22:43:05 -08:00
Kubernetes Prow Robot	06f0cba9b1	Merge pull request #115367 from tzneal/dedupe-resource-calculation dedupe pod resource request calculation	2023-03-09 22:42:50 -08:00
Kubernetes Prow Robot	1b647d5bf8	Merge pull request #114558 from TommyStarK/unit-tests/pkg-kubelet-nodestatus kubelet/nodestatus: Improving test coverage	2023-03-09 21:34:00 -08:00
Kubernetes Prow Robot	10802e9be1	Merge pull request #114498 from runzhliu/patch-2 Update kuberuntime_manager_test.go	2023-03-09 21:33:52 -08:00
Kubernetes Prow Robot	f6564d33ba	Merge pull request #114357 from dengyufeng2206/1208pull Log spelling formatting	2023-03-09 21:33:22 -08:00
Kubernetes Prow Robot	a3ad4d7623	Merge pull request #114017 from calvin0327/cleanup-containerruntime-options cleanup container runtime options	2023-03-09 21:33:06 -08:00
Kubernetes Prow Robot	15f5a5c6ef	Merge pull request #110949 from claudiubelu/adds-unittests-4 tests: Ports kubelet unit tests to Windows	2023-03-09 21:32:30 -08:00
Kubernetes Prow Robot	d241fcb4bd	Merge pull request #110760 from zhoumingcheng/master-unit-v2 add unit test coverage for pkg/kubelet/types/	2023-03-09 20:30:29 -08:00
Kubernetes Prow Robot	33d9543ceb	Merge pull request #111634 from KunWuLuan/pluginmanager_cache_log_amend docs(desired_state_of_world.go): log in desired_state_of_world.go seems to be wrong	2023-03-09 19:08:29 -08:00
Kubernetes Prow Robot	45b96eae98	Merge pull request #113145 from smarterclayton/zombie_terminating_pods kubelet: Force deleted pods can fail to move out of terminating	2023-03-09 15:32:30 -08:00
Todd Neal	4096c9209c	dedupe pod resource request calculation	2023-03-09 17:15:53 -06:00
Kubernetes Prow Robot	54ec651ab5	Merge pull request #110741 from zhoumingcheng/master-unit-v1 add unit test coverage for pkg/kubelet/util/queue	2023-03-09 11:15:51 -08:00
Clayton Coleman	6b9a381185	kubelet: Force deleted pods can fail to move out of terminating If a CRI error occurs during the terminating phase after a pod is force deleted (API or static) then the housekeeping loop will not deliver updates to the pod worker which prevents the pod's state machine from progressing. The pod will remain in the terminating phase but no further attempts to terminate or cleanup will occur until the kubelet is restarted. The pod worker now maintains a store of the pods state that it is attempting to reconcile and uses that to resync unknown pods when SyncKnownPods() is invoked, so that failures in sync methods for unknown pods no longer hang forever. The pod worker's store tracks desired updates and the last update applied on podSyncStatuses. Each goroutine now synchronizes to acquire the next work item, context, and whether the pod can start. This synchronization moves the pending update to the stored last update, which will ensure third parties accessing pod worker state don't see updates before the pod worker begins synchronizing them. As a consequence, the update channel becomes a simple notifier (struct{}) so that SyncKnownPods can coordinate with the pod worker to create a synthetic pending update for unknown pods (i.e. no one besides the pod worker has data about those pods). Otherwise the pending update info would be hidden inside the channel. In order to properly track pending updates, we have to be very careful not to mix RunningPods (which are calculated from the container runtime and are missing all spec info) and config- sourced pods. Update the pod worker to avoid using ToAPIPod() and instead require the pod worker to directly use update.Options.Pod or update.Options.RunningPod for the correct methods. Add a new SyncTerminatingRuntimePod to prevent accidental invocations of runtime only pod data. Finally, fix SyncKnownPods to replay the last valid update for undesired pods which drives the pod state machine towards termination, and alter HandlePodCleanups to: - terminate runtime pods that aren't known to the pod worker - launch admitted pods that aren't known to the pod worker Any started pods receive a replay until they reach the finished state, and then are removed from the pod worker. When a desired pod is detected as not being in the worker, the usual cause is that the pod was deleted and recreated with the same UID (almost always a static pod since API UID reuse is statistically unlikely). This simplifies the previous restartable pod support. We are careful to filter for active pods (those not already terminal or those which have been previously rejected by admission). We also force a refresh of the runtime cache to ensure we don't see an older version of the state. Future changes will allow other components that need to view the pod worker's actual state (not the desired state the podManager represents) to retrieve that info from the pod worker. Several bugs in pod lifecycle have been undetectable at runtime because the kubelet does not clearly describe the number of pods in use. To better report, add the following metrics: kubelet_desired_pods: Pods the pod manager sees kubelet_active_pods: "Admitted" pods that gate new pods kubelet_mirror_pods: Mirror pods the kubelet is tracking kubelet_working_pods: Breakdown of pods from the last sync in each phase, orphaned state, and static or not kubelet_restarted_pods_total: A counter for pods that saw a CREATE before the previous pod with the same UID was finished kubelet_orphaned_runtime_pods_total: A counter for pods detected at runtime that were not known to the kubelet. Will be populated at Kubelet startup and should never be incremented after. Add a metric check to our e2e tests that verifies the values are captured correctly during a serial test, and then verify them in detail in unit tests. Adds 23 series to the kubelet /metrics endpoint.	2023-03-08 22:03:51 -06:00
Paco Xu	a1def4b9c0	pod-infra-container-image: update comments as it will be removed in couple more releases Signed-off-by: Paco Xu <paco.xu@daocloud.io>	2023-03-09 11:14:32 +08:00
Kubernetes Prow Robot	625b8be09e	Merge pull request #115371 from pacoxu/cgroup-v2-memory-tuning default memoryThrottlingFactor to 0.9 and optimize the memory.high formulas	2023-03-08 18:46:00 -08:00
Kubernetes Prow Robot	8d5c96fed2	Merge pull request #116093 from swatisehgal/topologymanager-ga-graduation node: topologymgr: Graduate Kubelet Topology Manager to GA	2023-03-08 16:56:06 -08:00
calvin0327	0ffac50126	cleanup container runtime options Signed-off-by: calvin0327 <wen.chen@daocloud.io>	2023-03-08 16:53:19 +08:00
Paco Xu	f368413d65	sync default qps of kubelet change	2023-03-08 14:04:51 +08:00
Kubernetes Prow Robot	e390791e5f	Merge pull request #116341 from bobbypage/revert-114640-handle-device-mgr-recovery Revert "node: device-mgr: Handle recovery flow by checking if healthy devices exist"	2023-03-07 19:31:33 -08:00
Kubernetes Prow Robot	fe6a51ed4c	Merge pull request #116121 from wojtek-t/bump_qps_kubelet Bump default API QPS limits for Kubelet	2023-03-07 15:08:43 -08:00
Kubernetes Prow Robot	6bce018b36	Merge pull request #116271 from vinaykul/restart-free-pod-vertical-scaling-kubelet-panic-fix Fix nil pointer access panic in kubelet from uninitialized pod allocation checkpoint manager in standalone kubelet scenario	2023-03-07 12:38:45 -08:00
David Porter	9c20cee504	Revert "node: device-mgr: Handle recovery flow by checking if healthy devices exist"	2023-03-07 11:50:52 -08:00
Kubernetes Prow Robot	2c8f63f693	Merge pull request #115268 from jsafrane/split-reconstruction Split volume reconstruction refactoring from SELinuxMountReadWriteOncePod	2023-03-07 10:44:34 -08:00
Swati Sehgal	ae964a493f	node: topologymgr: remove comments with feature gate references Signed-off-by: Swati Sehgal <swsehgal@redhat.com>	2023-03-07 09:42:54 +00:00
vinay kulkarni	98e8f42f33	panic on pod resources alloc checkpoint failure	2023-03-07 05:59:34 +00:00
Kubernetes Prow Robot	8e659d43ec	Merge pull request #115925 from claudiubelu/skip-flaky-tests unit tests: Skip flaky tests on Windows	2023-03-06 21:56:29 -08:00
Kubernetes Prow Robot	44909771d9	Merge pull request #115965 from jsafrane/add-reconstruction-metrics Add volume reconstruction metrics	2023-03-06 14:56:16 -08:00
Claudiu Belu	5ba74c81ca	unit tests: Skip flaky tests on Windows Some of the unit tests are currently flaky on Windows. This commit skips them until they are resolved.	2023-03-06 20:46:05 +00:00
Jan Safranek	9ca548fcf0	Add metrics for force cleaned mounts after failed reconstruction Count nr. of force cleaned mounts + their failures after a volume fails reconstruction.	2023-03-06 17:48:59 +01:00
Kubernetes Prow Robot	d6e9cff212	Merge pull request #115838 from torredil/remove-aws Remove AWS legacy cloud provider + EBS in-tree storage plugin	2023-03-06 08:18:29 -08:00
Kubernetes Prow Robot	890d39f976	Merge pull request #114640 from swatisehgal/handle-device-mgr-recovery node: device-mgr: Handle recovery flow by checking if healthy devices exist	2023-03-06 07:10:28 -08:00
Kubernetes Prow Robot	68eea2468c	Merge pull request #114572 from huyinhou/fix-concurrent-map-access kubelet/deviceplugin: fix concurrent map iteration and map write	2023-03-06 06:06:29 -08:00
torredil	6aebda9b1e	Remove AWS legacy cloud provider + EBS in-tree storage plugin Signed-off-by: torredil <torredil@amazon.com>	2023-03-06 14:01:15 +00:00
Swati Sehgal	937d330393	node: topologymgr: Remove ResourceAllocator as TM is always enabled With Topology Manager enabled by default, we no longer need `resourceAllocator` as Topology Manager serves as the main PodAdmitHandler completely responsible for admission check based on hints received from the hintProviders and the subsequent allocation of the corresponding resources to a pod as can be seen here: https://github.com/kubernetes/kubernetes/blob/v1.26.0/pkg/kubelet/cm/topologymanager/scope.go#L150 With regard to DRA, the passing of `cm.draManager` into resourceAllocator seems redundant as no admission checks (and allocation of resources handled by DRA) is taking place in `Admit` method of resourceAllocator. DRA has a completely different model to the rest of the resource managers where pod is only scheduled on a node once resources are reserved for it. Because of this, admission checks or waiting for resources to be provisioned after the pod has been scheduled on the node is not required. Before making the above change, it was verified that DRA Manager is instantiated in `NewContainerManager`: https://github.com/kubernetes/kubernetes/blob/v1.26.0/pkg/kubelet/cm/container_manager_linux.go#L318 Signed-off-by: Swati Sehgal <swsehgal@redhat.com>	2023-03-06 12:51:11 +00:00
Swati Sehgal	6a62f0236a	node: topologymgr: trivial internal variable renaming Since Topology manager is graduating to GA, we remove internal configuration variable names with `Experimental` prefix. There is no expected change in behavior, only trival variable renaming. Signed-off-by: Swati Sehgal <swsehgal@redhat.com>	2023-03-06 12:51:11 +00:00
Swati Sehgal	d536a342b4	node: topologymgr: GA graduation implies Feature Gate is ON by default Signed-off-by: Swati Sehgal <swsehgal@redhat.com>	2023-03-06 12:51:05 +00:00
Swati Sehgal	5b2a3dbbdc	node: device-mgr: explicitly check if pre-allocated devices are healthy Signed-off-by: Swati Sehgal <swsehgal@redhat.com>	2023-03-06 11:52:23 +00:00
Swati Sehgal	a799ffb571	node: device-mgr: unit-tests: admission failure due to unhealthy devices Signed-off-by: Swati Sehgal <swsehgal@redhat.com>	2023-03-06 11:52:23 +00:00
Swati Sehgal	7ac399c205	node: device-mgr: Handle recovery by checking if healthy devices exist In case of node reboot/kubelet restart, the flow of events involves obtaining the state from the checkpoint file followed by setting the `healthDevices`/`unhealthyDevices` to its zero value. This is done to allow the device plugin to re-register itself so that capacity can be updated appropriately. During the allocation phase, we need to check if the resources requested by the pod have been registered AND healthy devices are present on the node to be allocated. Also we need to move this check above `needed==0` where needed is required - devices allocated to the container (which is obtained from the checkpoint file) because even in cases where no additional devices have to be allocated (as they were pre-allocated), we still need to make the devices that were previously allocated are healthy. Signed-off-by: Swati Sehgal <swsehgal@redhat.com>	2023-03-06 11:52:23 +00:00
Wojciech Tyczyński	280651abcc	Autogenerated	2023-03-06 12:08:34 +01:00
Wojciech Tyczyński	760acbbbe3	Bump QPS limits for Kubelet	2023-03-06 12:07:52 +01:00
SataQiu	528a471302	remove unused resize.go from pkg/kubelet/container	2023-03-06 18:33:13 +08:00
Kubernetes Prow Robot	b8aaaf380a	Merge pull request #116083 from SataQiu/clean-20230227 kubelet: remove unused DockerID type	2023-03-06 02:22:58 -08:00

1 2 3 4 5 ...

10539 Commits