1. Dec 11, 2020
    • Gerald Schaefer's avatar
      s390/mm: add support to allocate gigantic hugepages using CMA · 343dbdb7
      Gerald Schaefer authored
      Commit cf11e85f
      
       ("mm: hugetlb: optionally allocate gigantic hugepages
      using cma") added support for allocating gigantic hugepages using CMA,
      by specifying the hugetlb_cma= kernel parameter, which will disable any
      boot-time allocation of gigantic hugepages.
      
      This patch enables that option also for s390.
      
      Signed-off-by: default avatarGerald Schaefer <gerald.schaefer@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      343dbdb7
    • Harald Freudenberger's avatar
      s390/crypto: add arch_get_random_long() support · ff98cc98
      Harald Freudenberger authored
      
      
      The random longs to be pulled by arch_get_random_long() are
      prepared in an 4K buffer which is filled from the NIST 800-90
      compliant s390 drbg. By default the random long buffer is refilled
      256 times before the drbg itself needs a reseed. The reseed of the
      drbg is done with 32 bytes fetched from the high quality (but slow)
      trng which is assumed to deliver 100% entropy. So the 32 * 8 = 256
      bits of entropy are spread over 256 * 4KB = 1MB serving 131072
      arch_get_random_long() invocations before reseeded.
      
      How often the 4K random long buffer is refilled with the drbg
      before the drbg is reseeded can be adjusted. There is a module
      parameter 's390_arch_rnd_long_drbg_reseed' accessible via
        /sys/module/arch_random/parameters/rndlong_drbg_reseed
      or as kernel command line parameter
        arch_random.rndlong_drbg_reseed=<value>
      This parameter tells how often the drbg fills the 4K buffer before
      it is re-seeded by fresh entropy from the trng.
      A value of 16 results in reseeding the drbg at every 16 * 4 KB = 64
      KB with 32 bytes of fresh entropy pulled from the trng. So a value
      of 16 would result in 256 bits entropy per 64 KB.
      A value of 256 results in 1MB of drbg output before a reseed of the
      drbg is done. So this would spread the 256 bits of entropy among 1MB.
      Setting this parameter to 0 forces the reseed to take place every
      time the 4K buffer is depleted, so the entropy rises to 256 bits
      entropy per 4K or 0.5 bit entropy per arch_get_random_long().  With
      setting this parameter to negative values all this effort is
      disabled, arch_get_random long() returns false and thus indicating
      that the arch_get_random_long() feature is disabled at all.
      
      arch_get_random_long() is used by random.c among others to provide
      an initial hash value to be mixed with the entropy pool on every
      random data pull. For about 64 bytes read from /dev/urandom there
      is one call to arch_get_random_long(). So these additional random
      long values count for performance of /dev/urandom with measurable
      but low penalty.
      
      Signed-off-by: default avatarHarald Freudenberger <freude@linux.ibm.com>
      Reviewed-by: default avatarIngo Franzki <ifranzki@linux.ibm.com>
      Reviewed-by: default avatarJuergen Christ <jchrist@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      ff98cc98
  2. Dec 10, 2020
  3. Dec 03, 2020
  4. Nov 30, 2020
  5. Nov 23, 2020
    • Heiko Carstens's avatar
      s390/vdso: reimplement getcpu vdso syscall · 80f06306
      Heiko Carstens authored
      
      
      Implement the previously removed getcpu vdso syscall by using the
      TOD programmable field to pass the cpu number to user space.
      
      Reviewed-by: default avatarSven Schnelle <svens@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      80f06306
    • Heiko Carstens's avatar
      s390/mm: add debug user asce support · 062e5279
      Heiko Carstens authored
      
      
      Verify on exit to user space that always
      - the primary ASCE (cr1) is set to kernel ASCE
      - the secondary ASCE (cr7) is set to user ASCE
      
      If this is not the case: panic since something went terribly wrong.
      
      Reviewed-by: default avatarSven Schnelle <svens@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      062e5279
    • Heiko Carstens's avatar
      s390/mm: use invalid asce instead of kernel asce · 0290c9e3
      Heiko Carstens authored
      
      
      Create a region 3 page table which contains only invalid entries, and
      use that via "s390_invalid_asce" instead of the kernel ASCE whenever
      there is either
      - no user address space available, e.g. during early startup
      - as an intermediate ASCE when address spaces are switched
      
      This makes sure that user space accesses in such situations are
      guaranteed to fail.
      
      Reviewed-by: default avatarSven Schnelle <svens@linux.ibm.com>
      Reviewed-by: default avatarAlexander Gordeev <agordeev@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      0290c9e3
    • Heiko Carstens's avatar
      s390/mm: remove set_fs / rework address space handling · 87d59863
      Heiko Carstens authored
      
      
      Remove set_fs support from s390. With doing this rework address space
      handling and simplify it. As a result address spaces are now setup
      like this:
      
      CPU running in              | %cr1 ASCE | %cr7 ASCE | %cr13 ASCE
      ----------------------------|-----------|-----------|-----------
      user space                  |  user     |  user     |  kernel
      kernel, normal execution    |  kernel   |  user     |  kernel
      kernel, kvm guest execution |  gmap     |  user     |  kernel
      
      To achieve this the getcpu vdso syscall is removed in order to avoid
      secondary address mode and a separate vdso address space in for user
      space. The getcpu vdso syscall will be implemented differently with a
      subsequent patch.
      
      The kernel accesses user space always via secondary address space.
      This happens in different ways:
      - with mvcos in home space mode and directly read/write to secondary
        address space
      - with mvcs/mvcp in primary space mode and copy from primary space to
        secondary space or vice versa
      - with e.g. cs in secondary space mode and access secondary space
      
      Switching translation modes happens with sacf before and after
      instructions which access user space, like before.
      
      Lazy handling of control register reloading is removed in the hope to
      make everything simpler, but at the cost of making kernel entry and
      exit a bit slower. That is: on kernel entry the primary asce is always
      changed to contain the kernel asce, and on kernel exit the primary
      asce is changed again so it contains the user asce.
      
      In kernel mode there is only one exception to the primary asce: when
      kvm guests are executed the primary asce contains the gmap asce (which
      describes the guest address space). The primary asce is reset to
      kernel asce whenever kvm guest execution is interrupted, so that this
      doesn't has to be taken into account for any user space accesses.
      
      Reviewed-by: default avatarSven Schnelle <svens@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      87d59863
    • Heiko Carstens's avatar
      Merge branch 'fixes' into features · 77663819
      Heiko Carstens authored
      
      
      * fixes:
        s390: fix fpu restore in entry.S
      
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      77663819
    • Sven Schnelle's avatar
      s390: fix fpu restore in entry.S · 1179f170
      Sven Schnelle authored
      We need to disable interrupts in load_fpu_regs(). Otherwise an
      interrupt might come in after the registers are loaded, but before
      CIF_FPU is cleared in load_fpu_regs(). When the interrupt returns,
      CIF_FPU will be cleared and the registers will never be restored.
      
      The entry.S code usually saves the interrupt state in __SF_EMPTY on the
      stack when disabling/restoring interrupts. sie64a however saves the pointer
      to the sie control block in __SF_SIE_CONTROL, which references the same
      location.  This is non-obvious to the reader. To avoid thrashing the sie
      control block pointer in load_fpu_regs(), move the __SIE_* offsets eight
      bytes after __SF_EMPTY on the stack.
      
      Cc: <stable@vger.kernel.org> # 5.8
      Fixes: 0b0ed657
      
       ("s390: remove critical section cleanup from entry.S")
      Reported-by: default avatarPierre Morel <pmorel@linux.ibm.com>
      Signed-off-by: default avatarSven Schnelle <svens@linux.ibm.com>
      Acked-by: default avatarChristian Borntraeger <borntraeger@de.ibm.com>
      Reviewed-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      1179f170
  6. Nov 21, 2020
    • Heiko Carstens's avatar
      init/Kconfig: make COMPILE_TEST depend on !S390 · 334ef6ed
      Heiko Carstens authored
      While allmodconfig and allyesconfig build for s390 there are also
      various bots running compile tests with randconfig, where PCI is
      disabled. This reveals that a lot of drivers should actually depend on
      HAS_IOMEM.
      Adding this to each device driver would be a never ending story,
      therefore just disable COMPILE_TEST for s390.
      
      The reasoning is more or less the same as described in
      commit bc083a64
      
       ("init/Kconfig: make COMPILE_TEST depend on !UML").
      
      Reported-by: default avatarkernel test robot <lkp@intel.com>
      Suggested-by: default avatarArnd Bergmann <arnd@kernel.org>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      334ef6ed
    • Alexander Gordeev's avatar
      s390/vmem: make variable and function names consistent · 12bb4c68
      Alexander Gordeev authored
      
      
      Rename some variable and functions to better clarify
      what they are and what they do.
      
      Signed-off-by: default avatarAlexander Gordeev <agordeev@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      12bb4c68
    • Alexander Gordeev's avatar
    • Julian Wiedmann's avatar
      s390/stp: let subsys_system_register() sysfs attributes · 074ff04e
      Julian Wiedmann authored
      
      
      Instead of creating the sysfs attributes for the stp root_dev by hand,
      pass them to subsys_system_register() as parameter.
      
      This also ensures that the attributes are available when the KOBJ_ADD
      event is raised.
      
      Signed-off-by: default avatarJulian Wiedmann <jwi@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      074ff04e
    • Vasily Gorbik's avatar
      s390/decompressor: print cmdline and BEAR on pgm_check · ba1a6be9
      Vasily Gorbik authored
      
      
      Add kernel command line and last breaking event.
      The kernel command line is taken from early_command_line and printed
      only if kernel is not running as protected virtualization guest and
      if it has been already initialized from the "COMMAND_LINE".
      
      Linux version 5.10.0-rc3-22794-gecaa72788df0-dirty (gor@tuxmaker) #28 SMP PREEMPT Mon Nov 9 17:41:20 CET 2020
      Kernel command line: audit_enable=0 audit=0 selinux=0 crashkernel=296M root=/dev/dasda1 dasd=ec5b
      memblock=debug die
      Kernel fault: interruption code 0005 ilc:2
      PSW : 0000000180000000 0000000000012f92 (parse_boot_command_line+0x27a/0x46c)
            R:0 T:0 IO:0 EX:0 Key:0 M:0 W:0 P:0 AS:0 CC:0 PM:0 RI:0 EA:3
      GPRS: 0000000000000000 00ffffffffffffff 0000000000000000 000000000001a65c
            000000000000bf60 0000000000000000 00000000000003c0 0000000000000000
            0000000000000080 000000000002322d 000000007f29ef20 0000000000efd018
            000000000311c000 0000000000010070 0000000000012f82 000000000000bea8
      Call Trace:
      (sp:000000000000bea8 [<000000000002016e>] 000000000002016e)
       sp:000000000000bf18 [<0000000000012408>] startup_kernel+0x88/0x2fc
       sp:000000000000bf60 [<00000000000100c4>] startup_normal+0xb0/0xb0
      Last Breaking-Event-Address:
       [<00000000000135ba>] strcmp+0x22/0x24
      
      Reviewed-by: default avatarAlexander Egorenkov <egorenar@linux.ibm.com>
      Acked-by: default avatarViktor Mihajlovski <mihajlov@linux.ibm.com>
      Signed-off-by: default avatarVasily Gorbik <gor@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      ba1a6be9
    • Vasily Gorbik's avatar
      s390/decompressor: add stacktrace support · 8977ab65
      Vasily Gorbik authored
      
      
      Decompressor works on a single statically allocated stack. Stacktrace
      implementation with -mbackchain just takes few lines of code.
      
      Linux version 5.10.0-rc3-22793-g0f84a355b776-dirty (gor@tuxmaker) #27 SMP PREEMPT Mon Nov 9 17:30:18 CET 2020
      Kernel fault: interruption code 0005 ilc:2
      PSW : 0000000180000000 0000000000012f92 (parse_boot_command_line+0x27a/0x46c)
            R:0 T:0 IO:0 EX:0 Key:0 M:0 W:0 P:0 AS:0 CC:0 PM:0 RI:0 EA:3
      GPRS: 0000000000000000 00ffffffffffffff 0000000000000000 000000000001a62c
            000000000000bf60 0000000000000000 00000000000003c0 0000000000000000
            0000000000000080 000000000002322d 000000007f29ef20 0000000000efd018
            000000000311c000 0000000000010070 0000000000012f82 000000000000bea8
      Call Trace:
      (sp:000000000000bea8 [<000000000002016e>] 000000000002016e)
       sp:000000000000bf18 [<0000000000012408>] startup_kernel+0x88/0x2fc
       sp:000000000000bf60 [<00000000000100c4>] startup_normal+0xb0/0xb0
      
      Reviewed-by: default avatarAlexander Egorenkov <egorenar@linux.ibm.com>
      Signed-off-by: default avatarVasily Gorbik <gor@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      8977ab65
    • Vasily Gorbik's avatar
      s390/decompressor: add symbols support · 24621896
      Vasily Gorbik authored
      
      
      Information printed by print_pgm_check_info() is crucial for
      debugging decompressor problems. Printing instruction addresses is
      better than nothing, but turns further debugging into tedious job of
      figuring out which function those addresses correspond to.
      
      This change adds simplistic symbols resolution support. And adds %pS
      format specifier support to decompressor_printk().
      
      Decompressor symbols list is extracted and sorted with
      nm -n -S:
      ...
      0000000000010000 0000000000000014 T startup
      0000000000010014 00000000000000b0 t startup_normal
      0000000000010180 00000000000000b2 t startup_kdump
      ...
      
      Then functions are filtered and contracted to a form:
      "10000 14 startup\0""10014 b0 startup_normal\0""10180 b2 startup_kdump\0"
      ...
      Which makes it trivial to find beginning of an entry and names are 0
      terminated, so could be used as is. Symbols are binary-searched.
      
      To get symbols list with final addresses and then get it into the
      decompressor's image the same trick as for kallsyms is used.
      Decompressor's vmlinux is linked twice.
      
      Symbols are stored in .decompressor.syms section, current size is about
      2kb.
      
      Reviewed-by: default avatarAlexander Egorenkov <egorenar@linux.ibm.com>
      Signed-off-by: default avatarVasily Gorbik <gor@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      24621896
    • Vasily Gorbik's avatar
      s390/decompressor: correct some asm symbols annotations · ec55d1e1
      Vasily Gorbik authored
      
      
      Use SYM_CODE_* annotations for asm functions, so that function lengths
      are recognized correctly.
      
      Also currently the most part of startup is marked as startup_kdump. Move
      misplaced startup_kdump where it belongs.
      
      $ nm -n -S arch/s390/boot/compressed/vmlinux
      Before:
      0000000000010000 T startup
      0000000000010010 T startup_kdump
      After:
      0000000000010000 0000000000000014 T startup
      0000000000010014 00000000000000b0 t startup_normal
      0000000000010180 00000000000000b2 t startup_kdump
      
      Reviewed-by: default avatarAlexander Egorenkov <egorenar@linux.ibm.com>
      Signed-off-by: default avatarVasily Gorbik <gor@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      ec55d1e1
    • Vasily Gorbik's avatar
      s390/decompressor: add decompressor_printk · 9a78c70a
      Vasily Gorbik authored
      
      
      The decompressor does not have any special debug means. Running the
      kernel under qemu with gdb is helpful but tedious exercise if done
      repeatedly. It is also not applicable to debugging under LPAR and z/VM.
      
      One special thing which stands out is a working sclp_early_printk,
      which could be used once the kernel switches to 64-bit addressing mode.
      
      But sclp_early_printk does not provide any string formating capabilities.
      Formatting and printing string without printk-alike function is a
      not fun. The lack of printk-alike function means people would save up on
      testing and introduce more bugs.
      
      So, finally, provide decompressor_printk function, which fits on one
      screen and trades features for simplicity.
      
      It only supports "%s", "%x" and "%lx" specifiers and zero padding for
      hex values.
      
      Reviewed-by: default avatarAlexander Egorenkov <egorenar@linux.ibm.com>
      Signed-off-by: default avatarVasily Gorbik <gor@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      9a78c70a
    • Vasily Gorbik's avatar
      s390/ftrace: assume -mhotpatch or -mrecord-mcount always available · c9343637
      Vasily Gorbik authored
      
      
      Currently the kernel minimal compiler requirement is gcc 4.9 or
      clang 10.0.1.
      * gcc -mhotpatch option is supported since 4.8.
      * A combination of -pg -mrecord-mcount -mnop-mcount -mfentry flags is
      supported since gcc 9 and since clang 10.
      
      Drop support for old -pg function prologues. Which leaves binary
      compatible -mhotpatch / -mnop-mcount -mfentry prologues in a form:
      	brcl	0,0
      Which are also do not require initial nop optimization / conversion and
      presence of _mcount symbol.
      
      Signed-off-by: default avatarVasily Gorbik <gor@linux.ibm.com>
      Reviewed-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      c9343637
    • Vasily Gorbik's avatar
      s390: unify identity mapping limits handling · 73045a08
      Vasily Gorbik authored
      
      
      Currently we have to consider too many different values which
      in the end only affect identity mapping size. These are:
      1. max_physmem_end - end of physical memory online or standby.
         Always <= end of the last online memory block (get_mem_detect_end()).
      2. CONFIG_MAX_PHYSMEM_BITS - the maximum size of physical memory the
         kernel is able to support.
      3. "mem=" kernel command line option which limits physical memory usage.
      4. OLDMEM_BASE which is a kdump memory limit when the kernel is executed as
         crash kernel.
      5. "hsa" size which is a memory limit when the kernel is executed during
         zfcp/nvme dump.
      
      Through out kernel startup and run we juggle all those values at once
      but that does not bring any amusement, only confusion and complexity.
      
      Unify all those values to a single one we should really care, that is
      our identity mapping size.
      
      Signed-off-by: default avatarVasily Gorbik <gor@linux.ibm.com>
      Reviewed-by: default avatarAlexander Gordeev <agordeev@linux.ibm.com>
      Acked-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      73045a08
    • Julian Wiedmann's avatar
      s390/prng: let misc_register() add the prng sysfs attributes · 1e632eaa
      Julian Wiedmann authored
      
      
      Instead of creating the sysfs attributes for the prng devices by hand,
      describe them in .groups and let the misdevice core handle it.
      
      This also ensures that the attributes are available when the KOBJ_ADD
      event is raised.
      
      Signed-off-by: default avatarJulian Wiedmann <jwi@linux.ibm.com>
      Reviewed-by: default avatarHarald Freudenberger <freude@linux.ibm.com>
      Signed-off-by: default avatarHeiko Carstens <hca@linux.ibm.com>
      1e632eaa