1. Jan 24, 2023
    • Alexander Aring's avatar
      fs: dlm: move state change into else branch · ef7ef015
      Alexander Aring authored
      
      
      Currently we can switch at first into DLM_CLOSE_WAIT state and then do
      another state change if a condition is true. Instead of doing two state
      changes we handle the other state change inside an else branch of this
      condition.
      
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      ef7ef015
    • Alexander Aring's avatar
      fs: dlm: remove newline in log_print · 31864097
      Alexander Aring authored
      
      
      There is an API difference between log_print() and other printk()s to
      put a newline or not. This one was introduced by mistake because
      log_print() adds a newline.
      
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      31864097
    • Alexander Aring's avatar
      fs: dlm: reduce the shutdown timeout to 5 secs · 11605353
      Alexander Aring authored
      
      
      When a shutdown is stuck, time out after 5 seconds instead of
      3 minutes.  After this timeout we try a forced shutdown.
      
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      11605353
    • Alexander Aring's avatar
      fs: dlm: make dlm sequence id more robust · 317dd6ba
      Alexander Aring authored
      
      
      When joining a new lockspace, use a random number to initialize
      a sequence number used in messages. This makes it easier to detect
      sequence number mismatches in message replies during tests that
      repeatedly join and leave a lockspace.
      
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      317dd6ba
    • Alexander Aring's avatar
      fs: dlm: wait until all midcomms nodes detect version · b8b750e0
      Alexander Aring authored
      
      
      The current dlm version detection is very complex due to backwards
      compatablilty with earlier dlm protocol versions. It takes some time to
      detect if a peer node has a specific DLM version. If it's not detected,
      we just cut the socket connection. There could be cases where the local
      node has not detected the version yet, but the peer node has.  In these
      cases, we are trying to shutdown the dlm connection with a FIN/ACK message
      exchange to be sure the other peer is ready to shutdown the connection on
      dlm application level.  However this mechanism is only available on DLM
      protocol version 3.2 and we need to be sure the DLM version is detected
      before.
      
      To make it more robust we introduce a a "best effort" wait to wait for the
      version detection before shutdown the dlm connection. This need to be
      done before the kthread recoverd for recovery handling is stopped,
      because recovery handling will trigger enough messages to have a version
      detection going on.
      
      It is a corner case which was detected by modprobe dlm_locktroture module
      and rmmod dlm_locktorture module directly afterwards (in a looping
      behaviour). In practice probably nobody would leave a lockspace immediately
      after joining it.
      
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      b8b750e0
    • Alexander Aring's avatar
      fs: dlm: ignore unexpected non dlm opts msgs · 89835b06
      Alexander Aring authored
      
      
      This patch ignores unexpected RCOM_NAMES/RCOM_STATUS messages.
      To be backwards compatible, those messages are not part of the new
      reliable DLM OPTS encapsulation header, and have their own
      retransmit handling using sequence number matching  When we get
      unexpected non dlm opts messages, we should allow them and let
      RCOM message handling filter them out using sequence numbers.
      
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      89835b06
    • Alexander Aring's avatar
      fs: dlm: bring back previous shutdown handling · 54fbe0c1
      Alexander Aring authored
      This patch mostly reverts commit 4f567acb
      
       ("fs: dlm: remove socket
      shutdown handling"). There can be situations where the dlm midcomms nodes
      hash and lowcomms connection hash are not equal, but we need to guarantee
      that the lowcomms are all closed on a last release of a dlm lockspace,
      when a shutdown is invoked. This patch guarantees that we always close
      all sockets managed by the lowcomms connection hash, and calls shutdown
      for the last message sent. This ensures we don't cut the socket, which
      could cause the peer to get a connection reset.
      
      In future we should try to merge the midcomms/lowcomms hashes into one
      hash and not handle both in separate hashes.
      
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      54fbe0c1
    • Alexander Aring's avatar
      fs: dlm: send FIN ack back in right cases · 00908b33
      Alexander Aring authored
      This patch moves to send a ack back for receiving a FIN message only
      when we are in valid states. In other cases and there might be a sender
      waiting for a ack we just let it timeout at the senders time and
      hopefully all other cleanups will remove the FIN message on their
      sending queue. As an example we should never send out an ACK being in
      LAST_ACK state or we cannot assume a working socket communication when
      we are in CLOSED state.
      
      Cc: stable@vger.kernel.org
      Fixes: 489d8e55
      
       ("fs: dlm: add reliable connection if reconnect")
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      00908b33
    • Alexander Aring's avatar
      fs: dlm: move sending fin message into state change handling · a5849636
      Alexander Aring authored
      This patch moves the send fin handling, which should appear in a specific
      state change, into the state change handling while the per node
      state_lock is held. I experienced issues with other messages because
      we changed the state and a fin message was sent out in a different state.
      
      Cc: stable@vger.kernel.org
      Fixes: 489d8e55
      
       ("fs: dlm: add reliable connection if reconnect")
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      a5849636
    • Alexander Aring's avatar
      fs: dlm: don't set stop rx flag after node reset · 15c63db8
      Alexander Aring authored
      Similar to the stop tx flag, the rx flag should warn about a dlm message
      being received at DLM_FIN state change, when we are assuming no other
      dlm application messages. If we receive a FIN message and we are in the
      state DLM_FIN_WAIT2 we call midcomms_node_reset() which puts the
      midcomms node into DLM_CLOSED state. Afterwards we should not set the
      DLM_NODE_FLAG_STOP_RX flag any more.  This patch changes the setting
      DLM_NODE_FLAG_STOP_RX in those state changes when we receive a FIN
      message and we assume there will be no other dlm application messages
      received until we hit DLM_CLOSED state.
      
      Cc: stable@vger.kernel.org
      Fixes: 489d8e55
      
       ("fs: dlm: add reliable connection if reconnect")
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      15c63db8
    • Alexander Aring's avatar
      fs: dlm: fix race setting stop tx flag · 16427211
      Alexander Aring authored
      This patch sets the stop tx flag before we commit the dlm message.
      This flag will report about unexpected transmissions after we
      send the DLM_FIN message out, which should be the last message sent.
      When we commit the dlm fin message, it could be that we already
      got an ack back and the CLOSED state change already happened.
      We should not set this flag when we are in CLOSED state. To avoid this
      race we simply set the tx flag before the state change can be in
      progress by moving it before dlm_midcomms_commit_mhandle().
      
      Cc: stable@vger.kernel.org
      Fixes: 489d8e55
      
       ("fs: dlm: add reliable connection if reconnect")
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      16427211
    • Alexander Aring's avatar
      fs: dlm: be sure to call dlm_send_queue_flush() · 7354fa4e
      Alexander Aring authored
      If we release a midcomms node structure, there should be nothing left
      inside the dlm midcomms send queue. However, sometimes this is not true
      because I believe some DLM_FIN message was not acked... if we run
      into a shutdown timeout, then we should be sure there is no pending send
      dlm message inside this queue when releasing midcomms node structure.
      
      Cc: stable@vger.kernel.org
      Fixes: 489d8e55
      
       ("fs: dlm: add reliable connection if reconnect")
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      7354fa4e
    • Alexander Aring's avatar
      fs: dlm: fix use after free in midcomms commit · 724b6bab
      Alexander Aring authored
      While working on processing dlm message in softirq context I experienced
      the following KASAN use-after-free warning:
      
      [  151.760477] ==================================================================
      [  151.761803] BUG: KASAN: use-after-free in dlm_midcomms_commit_mhandle+0x19d/0x4b0
      [  151.763414] Read of size 4 at addr ffff88811a980c60 by task lock_torture/1347
      
      [  151.765284] CPU: 7 PID: 1347 Comm: lock_torture Not tainted 6.1.0-rc4+ #2828
      [  151.766778] Hardware name: Red Hat KVM/RHEL-AV, BIOS 1.16.0-3.module+el8.7.0+16134+e5908aa2 04/01/2014
      [  151.768726] Call Trace:
      [  151.769277]  <TASK>
      [  151.769748]  dump_stack_lvl+0x5b/0x86
      [  151.770556]  print_report+0x180/0x4c8
      [  151.771378]  ? kasan_complete_mode_report_info+0x7c/0x1e0
      [  151.772241]  ? dlm_midcomms_commit_mhandle+0x19d/0x4b0
      [  151.773069]  kasan_report+0x93/0x1a0
      [  151.773668]  ? dlm_midcomms_commit_mhandle+0x19d/0x4b0
      [  151.774514]  __asan_load4+0x7e/0xa0
      [  151.775089]  dlm_midcomms_commit_mhandle+0x19d/0x4b0
      [  151.775890]  ? create_message.isra.29.constprop.64+0x57/0xc0
      [  151.776770]  send_common+0x19f/0x1b0
      [  151.777342]  ? remove_from_waiters+0x60/0x60
      [  151.778017]  ? lock_downgrade+0x410/0x410
      [  151.778648]  ? __this_cpu_preempt_check+0x13/0x20
      [  151.779421]  ? rcu_lockdep_current_cpu_online+0x88/0xc0
      [  151.780292]  _convert_lock+0x46/0x150
      [  151.780893]  convert_lock+0x7b/0xc0
      [  151.781459]  dlm_lock+0x3ac/0x580
      [  151.781993]  ? 0xffffffffc0540000
      [  151.782522]  ? torture_stop+0x120/0x120 [dlm_locktorture]
      [  151.783379]  ? dlm_scan_rsbs+0xa70/0xa70
      [  151.784003]  ? preempt_count_sub+0xd6/0x130
      [  151.784661]  ? is_module_address+0x47/0x70
      [  151.785309]  ? torture_stop+0x120/0x120 [dlm_locktorture]
      [  151.786166]  ? 0xffffffffc0540000
      [  151.786693]  ? lockdep_init_map_type+0xc3/0x360
      [  151.787414]  ? 0xffffffffc0540000
      [  151.787947]  torture_dlm_lock_sync.isra.3+0xe9/0x150 [dlm_locktorture]
      [  151.789004]  ? torture_stop+0x120/0x120 [dlm_locktorture]
      [  151.789858]  ? 0xffffffffc0540000
      [  151.790392]  ? lock_torture_cleanup+0x20/0x20 [dlm_locktorture]
      [  151.791347]  ? delay_tsc+0x94/0xc0
      [  151.791898]  torture_ex_iter+0xc3/0xea [dlm_locktorture]
      [  151.792735]  ? torture_start+0x30/0x30 [dlm_locktorture]
      [  151.793606]  lock_torture+0x177/0x270 [dlm_locktorture]
      [  151.794448]  ? torture_dlm_lock_sync.isra.3+0x150/0x150 [dlm_locktorture]
      [  151.795539]  ? lock_torture_stats+0x80/0x80 [dlm_locktorture]
      [  151.796476]  ? do_raw_spin_lock+0x11e/0x1e0
      [  151.797152]  ? mark_held_locks+0x34/0xb0
      [  151.797784]  ? _raw_spin_unlock_irqrestore+0x30/0x70
      [  151.798581]  ? __kthread_parkme+0x79/0x110
      [  151.799246]  ? trace_preempt_on+0x2a/0xf0
      [  151.799902]  ? __kthread_parkme+0x79/0x110
      [  151.800579]  ? preempt_count_sub+0xd6/0x130
      [  151.801271]  ? __kasan_check_read+0x11/0x20
      [  151.801963]  ? __kthread_parkme+0xec/0x110
      [  151.802630]  ? lock_torture_stats+0x80/0x80 [dlm_locktorture]
      [  151.803569]  kthread+0x192/0x1d0
      [  151.804104]  ? kthread_complete_and_exit+0x30/0x30
      [  151.804881]  ret_from_fork+0x1f/0x30
      [  151.805480]  </TASK>
      
      [  151.806111] Allocated by task 1347:
      [  151.806681]  kasan_save_stack+0x26/0x50
      [  151.807308]  kasan_set_track+0x25/0x30
      [  151.807920]  kasan_save_alloc_info+0x1e/0x30
      [  151.808609]  __kasan_slab_alloc+0x63/0x80
      [  151.809263]  kmem_cache_alloc+0x1ad/0x830
      [  151.809916]  dlm_allocate_mhandle+0x17/0x20
      [  151.810590]  dlm_midcomms_get_mhandle+0x96/0x260
      [  151.811344]  _create_message+0x95/0x180
      [  151.811994]  create_message.isra.29.constprop.64+0x57/0xc0
      [  151.812880]  send_common+0x129/0x1b0
      [  151.813467]  _convert_lock+0x46/0x150
      [  151.814074]  convert_lock+0x7b/0xc0
      [  151.814648]  dlm_lock+0x3ac/0x580
      [  151.815199]  torture_dlm_lock_sync.isra.3+0xe9/0x150 [dlm_locktorture]
      [  151.816258]  torture_ex_iter+0xc3/0xea [dlm_locktorture]
      [  151.817129]  lock_torture+0x177/0x270 [dlm_locktorture]
      [  151.817986]  kthread+0x192/0x1d0
      [  151.818518]  ret_from_fork+0x1f/0x30
      
      [  151.819369] Freed by task 1336:
      [  151.819890]  kasan_save_stack+0x26/0x50
      [  151.820514]  kasan_set_track+0x25/0x30
      [  151.821128]  kasan_save_free_info+0x2e/0x50
      [  151.821812]  __kasan_slab_free+0x107/0x1a0
      [  151.822483]  kmem_cache_free+0x204/0x5e0
      [  151.823152]  dlm_free_mhandle+0x18/0x20
      [  151.823781]  dlm_mhandle_release+0x2e/0x40
      [  151.824454]  rcu_core+0x583/0x1330
      [  151.825047]  rcu_core_si+0xe/0x20
      [  151.825594]  __do_softirq+0xf4/0x5c2
      
      [  151.826450] Last potentially related work creation:
      [  151.827238]  kasan_save_stack+0x26/0x50
      [  151.827870]  __kasan_record_aux_stack+0xa2/0xc0
      [  151.828609]  kasan_record_aux_stack_noalloc+0xb/0x20
      [  151.829415]  call_rcu+0x4c/0x760
      [  151.829954]  dlm_mhandle_delete+0x97/0xb0
      [  151.830718]  dlm_process_incoming_buffer+0x2fc/0xb30
      [  151.831524]  process_dlm_messages+0x16e/0x470
      [  151.832245]  process_one_work+0x505/0xa10
      [  151.832905]  worker_thread+0x67/0x650
      [  151.833507]  kthread+0x192/0x1d0
      [  151.834046]  ret_from_fork+0x1f/0x30
      
      [  151.834900] The buggy address belongs to the object at ffff88811a980c30
                      which belongs to the cache dlm_mhandle of size 88
      [  151.836894] The buggy address is located 48 bytes inside of
                      88-byte region [ffff88811a980c30, ffff88811a980c88)
      
      [  151.839007] The buggy address belongs to the physical page:
      [  151.839904] page:0000000076cf5d62 refcount:1 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x11a980
      [  151.841378] flags: 0x8000000000000200(slab|zone=2)
      [  151.842141] raw: 8000000000000200 0000000000000000 dead000000000122 ffff8881089b43c0
      [  151.843401] raw: 0000000000000000 0000000000220022 00000001ffffffff 0000000000000000
      [  151.844640] page dumped because: kasan: bad access detected
      
      [  151.845822] Memory state around the buggy address:
      [  151.846602]  ffff88811a980b00: fb fb fb fb fc fc fc fc fa fb fb fb fb fb fb fb
      [  151.847761]  ffff88811a980b80: fb fb fb fc fc fc fc fa fb fb fb fb fb fb fb fb
      [  151.848921] >ffff88811a980c00: fb fb fc fc fc fc fa fb fb fb fb fb fb fb fb fb
      [  151.850076]                                                        ^
      [  151.851085]  ffff88811a980c80: fb fc fc fc fc fa fb fb fb fb fb fb fb fb fb fb
      [  151.852269]  ffff88811a980d00: fc fc fc fc fa fb fb fb fb fb fb fb fb fb fb fc
      [  151.853428] ==================================================================
      [  151.855618] Disabling lock debugging due to kernel taint
      
      It is accessing a mhandle in dlm_midcomms_commit_mhandle() and the mhandle
      was freed by a call_rcu() call in dlm_process_incoming_buffer(),
      dlm_mhandle_delete(). It looks like it was freed because an ack of
      this message was received. There is a short race between committing the
      dlm message to be transmitted and getting an ack back. If the ack is
      faster than returning from dlm_midcomms_commit_msg_3_2(), then we run
      into a use-after free because we still need to reference the mhandle when
      calling srcu_read_unlock().
      
      To avoid that, we don't allow that mhandle to be freed between
      dlm_midcomms_commit_msg_3_2() and srcu_read_unlock() by using rcu read
      lock. We can do that because mhandle is protected by rcu handling.
      
      Cc: stable@vger.kernel.org
      Fixes: 489d8e55
      
       ("fs: dlm: add reliable connection if reconnect")
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      724b6bab
    • Alexander Aring's avatar
      fs: dlm: start midcomms before scand · aad633dc
      Alexander Aring authored
      The scand kthread can send dlm messages out, especially dlm remove
      messages to free memory for unused rsb on other nodes. To send out dlm
      messages, midcomms must be initialized. This patch moves the midcomms
      start before scand is started.
      
      Cc: stable@vger.kernel.org
      Fixes: e7fd4179
      
       ("[DLM] The core of the DLM for GFS2/CLVM")
      Signed-off-by: default avatarAlexander Aring <aahringo@redhat.com>
      Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
      aad633dc
  2. Jan 06, 2023
  3. Jan 02, 2023
  4. Jan 01, 2023
  5. Dec 31, 2022
    • Linus Torvalds's avatar
      Merge tag 'acpi-6.2-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm · c8451c14
      Linus Torvalds authored
      Pull ACPI fixes from Rafael Wysocki:
       "These are new ACPI IRQ override quirks, low-power S0 idle (S0ix)
        support adjustments and ACPI backlight handling fixes, mostly for
        platforms using AMD chips.
      
        Specifics:
      
         - Add ACPI IRQ override quirks for Asus ExpertBook B2502, Lenovo
           14ALC7, and XMG Core 15 (Hans de Goede, Adrian Freund, Erik
           Schumacher).
      
         - Adjust ACPI video detection fallback path to prevent
           non-operational ACPI backlight devices from being created on
           systems where the native driver does not detect a suitable panel
           (Mario Limonciello).
      
         - Fix Apple GMUX backlight detection (Hans de Goede).
      
         - Add a low-power S0 idle (S0ix) handling quirk for HP Elitebook 865
           and stop using AMD-specific low-power S0 idle code path for systems
           with Rembrandt chips and newer (Mario Limonciello)"
      
      * tag 'acpi-6.2-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm:
        ACPI: x86: s2idle: Stop using AMD specific codepath for Rembrandt+
        ACPI: x86: s2idle: Force AMD GUID/_REV 2 on HP Elitebook 865
        ACPI: video: Fix Apple GMUX backlight detection
        ACPI: resource: Add Asus ExpertBook B2502 to Asus quirks
        ACPI: resource: do IRQ override on Lenovo 14ALC7
        ACPI: resource: do IRQ override on XMG Core 15
        ACPI: video: Don't enable fallback path for creating ACPI backlight by default
        drm/amd/display: Report to ACPI video if no panels were found
        ACPI: video: Allow GPU drivers to report no panels
      c8451c14
    • Linus Torvalds's avatar
      Merge tag 'sound-6.2-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound · 262eef26
      Linus Torvalds authored
      Pull sound fixes from Takashi Iwai:
       "Just a few small fixes:
      
         - A regression fix for HDMI audio on HD-audio AMD codecs
      
         - Fixes for LINE6 MIDI handling
      
         - HD-audio quirk for Dell laptops"
      
      * tag 'sound-6.2-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound:
        ALSA: hda/hdmi: Static PCM mapping again with AMD HDMI codecs
        ALSA: hda/realtek: Apply dual codec fixup for Dell Latitude laptops
        ALSA: line6: fix stack overflow in line6_midi_transmit
        ALSA: line6: correct midi status byte when receiving data from podxt
      262eef26
  6. Dec 30, 2022