kubernetes

Author	SHA1	Message	Date
Kevin Klues	708278098a	Update bitmask printing to print in groups of 2 instead of all 64 bits	2020-01-16 17:28:52 +01:00
Kevin Klues	7069b1d6e8	Update TopologyManager single-numa-node logic to handle "don't cares" The logic has been updated to match the logic of the best-effort policy except in two places: 1) The hint filtering frunction has been updated to allow "don't care" hints encoded with a `nil` affinity mask, to pass through the filter in addition to hints that have just a single NUMA bit set. 2) After calculating the `bestHint` we transform "don't care" affinities encoded as having all NUMA bits set in their affinity masks into "don't care" affinities encoded as `nil`.	2020-01-16 08:50:35 +00:00
Kevin Klues	2905ffffa7	Rename TopologyManager test TestPolicyBestEffortMerge for consistency	2020-01-16 08:50:21 +00:00
Kevin Klues	94489c137c	Cleanup use of defaultAffinity in mergePermutation of TopologyManager	2020-01-16 08:50:12 +00:00
nolancon	5e23517ebf	Use reflect.DeepEqual check in policy_test.go	2020-01-16 08:13:07 +00:00
nolancon	92eb7cd601	Update "Single NUMA hint generation" expected affinity to nil	2020-01-16 08:13:07 +00:00
nolancon	8b3f6e61a2	Move test case "Two providers, 1 with 2 hints, 1 with single non-preferred hint matching" into specific policy tests	2020-01-16 08:13:07 +00:00
nolancon	681c42bfc2	Move test case "Two providers, 1 hint each, same mask, 1 preferred, 1 not 2/2" into specific policy tests	2020-01-16 08:13:07 +00:00
nolancon	a38a2562b2	Move test case "Two providers, 1 hint each, same mask, 1 preferred, 1 not 1/2" into specific policy test.	2020-01-16 08:13:07 +00:00
nolancon	f639da7637	Move test case "Two providers, 1 hint each, no common mask" into specific policy tests.	2020-01-16 08:13:07 +00:00
nolancon	401a2bb285	Move test case "Single TopologyHint with Preferred as false and NUMANodeAffinity as nil" into specific policy tests.	2020-01-16 08:13:06 +00:00
nolancon	6460ef6392	Move test case "Single TopologyHint with Preferred as true and NUMANodeAffinity as nil" into specific policy tests.	2020-01-16 08:13:06 +00:00
nolancon	baeff9ec5d	Move test case "HintProvider returns empty non-nil map[string][]TopologyHint from provider" into specific policy tests.	2020-01-16 08:13:06 +00:00
nolancon	599217d482	Move test case "HintProvider returns -nil map[string][]TopologyHint from provider" into specific policy tests	2020-01-16 08:13:06 +00:00
nolancon	57661ee946	Move test case 'HintProvider returns empty non-nil map[string][]TopologyHint' into specific policy tests.	2020-01-16 08:13:06 +00:00
nolancon	51f1af0395	Move test case 'TopologyHint not set' into individual policy tests	2020-01-16 08:13:06 +00:00
nolancon	8466a5852a	Restore policy_test.go to upstream Following commits will contain incremental changes to this file to ease review process and ensure all tests are accounted for.	2020-01-16 08:13:06 +00:00
nolancon	59bb6c4d6f	Update checks in mergeProvidersHints: - Initialize best Hint to TopologyHint{} - Update checks. - Move generic unit test case into policy specific tests and updated expected outcome to reflect changes.	2020-01-16 08:13:06 +00:00
nolancon	6758f95117	Restore original policy none test cases: Mistakenly overwritten in earlier commit	2020-01-16 08:13:06 +00:00
nolancon	2d1a535a35	Make mergePermutation generic: - Remove policy parameters to make function generic - Move function into top level policy.go	2020-01-16 08:13:06 +00:00
nolancon	5487941485	Refactor filterHints: - Restructure function - Remove bug fix for catching {nil true} - To be fixed in later commit - Restore unit tests to original state for testing filterHints	2020-01-16 08:13:06 +00:00
nolancon	adfd11f38f	Make iterateAllProviderTopologyHints generic: - Remove policy parameters to make this function generic. - Move function out of individual policies and into policy.go	2020-01-16 08:13:06 +00:00
nolancon	e43f0a5293	Reinstate canAdmitPodResult in policy_none: This is to keep consistency with the other policies. This change may be made across all policies in a future PR, but removing it from the scope of this PR for now.	2020-01-16 08:13:05 +00:00
nolancon	4cc5b9e46c	Edit hints returned from policies and unit tests: - Best Effort Policy: Return hint with nil affinity as opposed to defaultAffinity when provider has no preference for NUMA affinty or no possible NUMA affinities. - Single NUMA Node Policy: Remove defaultHint from mergeProvidersHints. Instead return appropriate TopologyHint where required. - Update unit tests to reflect changes. Some test cases moved into individual policy test functions due to differing returned affinties per policy.	2020-01-16 08:13:05 +00:00
nolancon	e3d0c9397f	Updates to single-numa-node policy: - Remove getHintMatch method. - Replace with simplified versions of mergePermutation and iterateAllProviderTopologyHints methods - as used in best-effort. - Remove getHintMatch unit tests.	2020-01-16 08:13:05 +00:00
nolancon	b5ca4989e3	Update unit tests: - Update filterHints test to reflect changes in previous commit. - Some common test cases achieve differing expected results based on policy due to independent merge strategies. These cases are moved into individual policy based test functions.	2020-01-16 08:13:05 +00:00
nolancon	17d615bca2	Update filterHints: - Only append valid preferred-true hints to filtered - Return true if allResourceHints only consist of nil-affinity/preferred-true hints: {nil true}, update defaultHint preference accordingly.	2020-01-16 08:13:05 +00:00
Adrian Chiris	9f21f49493	Additional unit tests for Topology Manager methods	2020-01-16 08:13:05 +00:00
Adrian Chiris	f886d2a832	Update single-numa-node policy unit tests	2020-01-16 08:13:05 +00:00
Adrian Chiris	2825a7be1a	Add new functionality for single-numa-node policy: Explanation taken from original commit: - Change the current method of finding the best hint. Instead of going over all permutations, sort the hints and find the narrowest hint common to all resources. - Break out early when merging to a preferred hint is not possible	2020-01-16 08:13:05 +00:00
Adrian Chiris	5ce2ea2773	Return defaultAffinity from PolicyBestEffort: Now that PolicySingleNUMANode is not considered here, return defaultAffinity as was the original case before previous bug fix	2020-01-16 08:13:05 +00:00
Adrian Chiris	eda1521562	Make mergeProviderHints policy-specific: - Remove need to pass policy and numaNodes as arguments - Remove PolicySingleNUMANode special case check in policy_best_effort - Add mergeProviderHints base to policy_single_numa_node for upcoming commit	2020-01-16 08:13:05 +00:00
Adrian Chiris	dc36924c37	Update policy_none removing canAdmitPodResult Update unit tests for none_policy Add Name test for policy_restricted	2020-01-16 08:13:05 +00:00
Adrian Chiris	cf8b098dda	Refactor policy-best-effort - Modularize code with mergePermutation method	2020-01-16 08:13:05 +00:00
Sascha Grunert	278717bc57	Fix ineffectual assignment to CPUSets Signed-off-by: Sascha Grunert <sgrunert@suse.com>	2020-01-16 08:57:42 +01:00
Kevin Klues	34b942a41d	Remove check for empty activePods list in CPUManager removeStaleState This check is redundant since we protect this call with a call to `m.sourcesReady.AllReady()` earlier on. Moreover, having this check in place means that we will leave some stale state around in cases where there are actually no active pods in the system and this loop hasn't cleaned them up yet. This can happen, for example, if a pod exits while the kubelet is down for some reason. We see this exact case being triggered in our e2e tests, where a test has been failing since October when this change was first introduced.	2020-01-15 20:09:24 +01:00
Kevin Klues	5802f3a910	Add proper activePods list in TestGetTopologyHints for CPUManager	2020-01-15 20:08:41 +01:00
danielqsj	1a9b121764	remove deprecated metrics of kubelet	2020-01-10 16:46:52 +08:00
Kubernetes Prow Robot	fd0358fd21	Merge pull request #86689 from klueska/upstream-fix-cpumanager-v1-state-checksum Lock checksum calculation for v1 CPUManager state to pre 1.18 logic	2020-01-08 02:57:40 -08:00
Kubernetes Prow Robot	d6412b856f	Merge pull request #84345 from danielqsj/withdialer replace grpc.WithDialer which is deprecated	2020-01-06 15:56:17 -08:00
Kubernetes Prow Robot	9acf7d11fe	Merge pull request #86344 from klueska/upstream-cm-approver Add klueska as an approver in pkg/kubelet/cm/OWNERS	2020-01-06 09:54:16 -08:00
Kevin Klues	b373121a14	Make CPUManagerCheckpointV2 type an alias of CPUManagerCheckpoint This change is to prevent problems when we remove the V1->V2 migration code in the future. Without this, the checksums of all checkpoints would be hashed with the name CPUManagerCheckpointV2 embedded inside of them, which is undesirable. We want the checkpoints to be hashed with the name CPUManagerCheckpoint instead.	2019-12-28 19:29:13 +01:00
Kevin Klues	5faf8f4c52	Lock checksum calculation for v1 CPUManager state to pre 1.18 logic The updated CPUManager from PR #84462 implements logic to migrate the CPUManager checkpoint file from an old format to a new one. To do so, it defines the following types: ``` type CPUManagerCheckpoint = CPUManagerCheckpointV2 type CPUManagerCheckpointV1 struct { ... } type CPUManagerCheckpointV2 struct { ... } ``` This replaces the old definition of just: ``` type CPUManagerCheckpoint struct { ... } ``` Code was put in place to ensure proper migration from checkpoints in V1 format to checkpoints in V2 format. However (and this is a big however), all of the unit tests were performed on V1 checkpoints that were generated using the type name `CPUManagerCheckpointV1` and not the original type name of `CPUManagerCheckpoint`. As such, the checksum in the checkpoint file uses the `CPUManagerCheckpointV1` type to calculate its checksum and not the original type name of `CPUManagerCheckpoint`. This causes problems in the real world since all pre-1.18 checkpoint files will have been generated with the original type name of `CPUManagerCheckpoint`. When verifying the checksum of the checkpoint file across an upgrade to 1.18, the checksum is calculated assuming a type name of `CPUManagerCheckpointV1` (which is incorrect) and the file is seen to be corrupt. This patch ensures that all V1 checksums are verified against a type name of `CPUManagerCheckpoint` instead of ``CPUManagerCheckpointV1`. It also locks the algorithm used to calculate the checksum in place, since it wil never change in the future (for pre-1.18 checkpoint files at least).	2019-12-28 14:17:55 +01:00
danielqsj	19fe9f8d94	replace grpc.WithDialer which is deprecated	2019-12-26 17:46:59 +08:00
whypro	f4bd4e2e96	Return error instead of panic when cpu manager starts failed.	2019-12-19 21:56:23 +08:00
Kevin Klues	9818b4522e	Add klueska as an approver in pkg/kubelet/cm/OWNERS	2019-12-17 10:40:23 +01:00
Kevin Klues	f553286156	Pass initial set of runtime containers to the CPUManager at startup These information associatedd with these containers is used to migrate the CPUManager state from it's old format to its new (i.e. keyed off of podUID and containerName instead of containerID).	2019-12-11 23:02:51 +01:00
Kevin Klues	6441e1ef43	Move CPUManager Checkpoint restoration to Start() instead of New()	2019-12-11 23:02:51 +01:00
Kevin Klues	69f8053850	Update top-level CPUManager to adhere to new state semantics For now, we just pass 'nil' as the set of 'initialContainers' for migrating from old state semantics to new ones. In a subsequent commit will we pull this information from higher layers so that we can pass it down at this stage properly.	2019-12-11 23:02:51 +01:00
Kevin Klues	185e790f71	Update CPUManager policies to adhere to new state semantics	2019-12-11 23:02:51 +01:00

1 2 3 4 5 ...

791 Commits