Handle non-ASCII identifiers in Ada

Ada allows non-ASCII identifiers, and GNAT supports several such encodings. This patch adds the corresponding support to gdb. GNAT encodes non-ASCII characters using special symbol names. For character sets like Latin-1, where all characters are a single byte, it uses a "U" followed by the hex for the character. So, for example, thorn would be encoded as "Ufe" (0xFE being lower case thorn). For wider characters, despite what the manual says (it claims Shift-JIS and EUC can be used), in practice recent versions only support Unicode. Here, characters in the base plane are represented using "Wxxxx" and characters outside the base plane using "WWxxxxxxxx". GNAT has some further quirks here. Ada is case-insensitive, and GNAT emits symbols that have been case-folded. For characters in ASCII, and for all characters in non-Unicode character sets, lower case is used. For Unicode, however, characters that fit in a single byte are converted to lower case, but all others are converted to upper case. Furthermore, there is a bug in GNAT where two symbols that differ only in the case of "Y WITH DIAERESIS" (and potentially others, I did not check exhaustively) can be used in one program. I chose to omit handling this case from gdb, on the theory that it is hard to figure out the logic, and anyway if the bug is ever fixed, we'll regret having a heuristic. This patch introduces a new "ada source-charset" setting. It defaults to Latin-1, as that is GNAT's default. This setting controls how "U" characters are decoded -- W/WW are always handled as UTF-32. The ada_tag_name_from_tsd change is needed because this function will read memory from the inferior and interpret it -- and this caused an encoding failure on PPC when running a test that tries to read uninitialized memory. This patch implements its own UTF-32-based case folder. This avoids host platform quirks, and is relatively simple. A short Python program to generate the case-folding table is included. It simply relies on whatever version of Unicode is used by the host Python, which seems basically acceptable. Test cases for UTF-8, Latin-1, and Latin-3 are included. This exercises most of the new code paths, aside from Y WITH DIAERESIS as noted above.
author: Tom Tromey <tromey@adacore.com> 2022-02-03 10:42:07 -0700
committer: Tom Tromey <tromey@adacore.com> 2022-03-07 07:52:59 -0700
commit: 315e4ebb4b7ef01da2f5c419edc74f39a0122d20 (patch)
tree: ed8a010b58b1f7cb532b83d602b39adfc07397f8 /gdb/doc
parent: ee3d46491537e343c276a7fc455dd94812fd3f72 (diff)
download: gdb-315e4ebb4b7ef01da2f5c419edc74f39a0122d20.zip
gdb-315e4ebb4b7ef01da2f5c419edc74f39a0122d20.tar.gz
gdb-315e4ebb4b7ef01da2f5c419edc74f39a0122d20.tar.bz2
1 files changed, 23 insertions, 0 deletions
diff --git a/gdb/doc/gdb.texinfo b/gdb/doc/gdb.texinfo
index a80b61e..5d1dcfd 100644
--- a/gdb/doc/gdb.texinfo
+++ b/gdb/doc/gdb.texinfo
@@ -18012,6 +18012,7 @@ to be difficult.
 * Ravenscar Profile::           Tasking Support when using the Ravenscar
                                    Profile
 * Ada Settings::                New settable GDB parameters for Ada.
+* Ada Source Character Set::    Character set of Ada source files.
 * Ada Glitches::                Known peculiarities of Ada mode.
 @end menu
 
@@ -18762,6 +18763,28 @@ size is less than @var{size}.
 Show the limit on types whose size is determined by run-time quantities.
 @end table
 
+@node Ada Source Character Set
+@subsubsection Ada Source Character Set
+@cindex Ada, source character set
+
+The GNAT compiler supports a number of character sets for source
+files.  @xref{Character Set Control, , Character Set Control,
+gnat_ugn}.  @value{GDBN} includes support for this as well.
+
+@table @code
+@item set ada source-charset @var{charset}
+@kindex set ada source-charset
+Set the source character set for Ada.  The character set must be
+supported by GNAT.  Because this setting affects the decoding of
+symbols coming from the debug information in your program, the setting
+should be set as early as possible.  The default is @code{ISO-8859-1},
+because that is also GNAT's default.
+
+@item show ada source-charset
+@kindex show ada source-charset
+Show the current source character set for Ada.
+@end table
+
 @node Ada Glitches
 @subsubsection Known Peculiarities of Ada Mode
 @cindex Ada, problems
author	Tom Tromey <tromey@adacore.com>	2022-02-03 10:42:07 -0700
committer	Tom Tromey <tromey@adacore.com>	2022-03-07 07:52:59 -0700
commit	315e4ebb4b7ef01da2f5c419edc74f39a0122d20 (patch)
tree	ed8a010b58b1f7cb532b83d602b39adfc07397f8 /gdb/doc
parent	ee3d46491537e343c276a7fc455dd94812fd3f72 (diff)
download	gdb-315e4ebb4b7ef01da2f5c419edc74f39a0122d20.zip gdb-315e4ebb4b7ef01da2f5c419edc74f39a0122d20.tar.gz gdb-315e4ebb4b7ef01da2f5c419edc74f39a0122d20.tar.bz2