diff options
author | Lewis Hyatt <lhyatt@gmail.com> | 2019-12-09 20:03:47 +0000 |
---|---|---|
committer | David Malcolm <dmalcolm@gcc.gnu.org> | 2019-12-09 20:03:47 +0000 |
commit | ee9256409f21eab5df5076e46d220d6a0b995f79 (patch) | |
tree | 68053762905d3e64e86dc0db19b7d1f3d65b5ba8 /contrib/unicode/README | |
parent | 763c9f4a8544318998c7adf04e4c92e9a4b85614 (diff) | |
download | gcc-ee9256409f21eab5df5076e46d220d6a0b995f79.zip gcc-ee9256409f21eab5df5076e46d220d6a0b995f79.tar.gz gcc-ee9256409f21eab5df5076e46d220d6a0b995f79.tar.bz2 |
Byte vs column awareness for diagnostic-show-locus.c (PR 49973)
contrib/ChangeLog
2019-12-09 Lewis Hyatt <lhyatt@gmail.com>
PR preprocessor/49973
* unicode/from_glibc/unicode_utils.py: Support script from
glibc (commit 464cd3) to extract character widths from Unicode data
files.
* unicode/from_glibc/utf8_gen.py: Likewise.
* unicode/UnicodeData.txt: Unicode v. 12.1.0 data file.
* unicode/EastAsianWidth.txt: Likewise.
* unicode/PropList.txt: Likewise.
* unicode/gen_wcwidth.py: New utility to generate
libcpp/generated_cpp_wcwidth.h with help from the glibc support
scripts and the Unicode data files.
* unicode/unicode-license.txt: Added.
* unicode/README: New explanatory file.
libcpp/ChangeLog
2019-12-09 Lewis Hyatt <lhyatt@gmail.com>
PR preprocessor/49973
* generated_cpp_wcwidth.h: New file generated by
../contrib/unicode/gen_wcwidth.py, supports new cpp_wcwidth function.
* charset.c (compute_next_display_width): New function to help
implement display columns.
(cpp_byte_column_to_display_column): Likewise.
(cpp_display_column_to_byte_column): Likewise.
(cpp_wcwidth): Likewise.
* include/cpplib.h (cpp_byte_column_to_display_column): Declare.
(cpp_display_column_to_byte_column): Declare.
(cpp_wcwidth): Declare.
(cpp_display_width): New function.
gcc/ChangeLog
2019-12-09 Lewis Hyatt <lhyatt@gmail.com>
PR preprocessor/49973
* input.c (location_compute_display_column): New function to help with
multibyte awareness in diagnostics.
(test_cpp_utf8): New self-test.
(input_c_tests): Call the new test.
* input.h (location_compute_display_column): Declare.
* diagnostic-show-locus.c: Pervasive changes to add multibyte awareness
to all classes and functions.
(enum column_unit): New enum.
(class exploc_with_display_col): New class.
(class layout_point): Convert m_column member to array m_columns[2].
(layout_range::contains_point): Add col_unit argument.
(test_layout_range_for_single_point): Pass new argument.
(test_layout_range_for_single_line): Likewise.
(test_layout_range_for_multiple_lines): Likewise.
(line_bounds::convert_to_display_cols): New function.
(layout::get_state_at_point): Add col_unit argument.
(make_range): Use empty filename rather than dummy filename.
(get_line_width_without_trailing_whitespace): Rename to...
(get_line_bytes_without_trailing_whitespace): ...this.
(test_get_line_width_without_trailing_whitespace): Rename to...
(test_get_line_bytes_without_trailing_whitespace): ...this.
(class layout): m_exploc changed to exploc_with_display_col from
plain expanded_location.
(layout::get_linenum_width): New accessor member function.
(layout::get_x_offset_display): Likewise.
(layout::calculate_linenum_width): New subroutine for the constuctor.
(layout::calculate_x_offset_display): Likewise.
(layout::layout): Use the new subroutines. Add multibyte awareness.
(layout::print_source_line): Add multibyte awareness.
(layout::print_line): Likewise.
(layout::print_annotation_line): Likewise.
(line_label::line_label): Likewise.
(layout::print_any_labels): Likewise.
(layout::annotation_line_showed_range_p): Likewise.
(get_printed_columns): Likewise.
(class line_label): Rename m_length to m_display_width.
(get_affected_columns): Rename to...
(get_affected_range): ...this; add col_unit argument and multibyte
awareness.
(class correction): Add m_affected_bytes and m_display_cols
members. Rename m_len to m_byte_length for clarity. Add multibyte
awareness throughout.
(correction::insertion_p): Add multibyte awareness.
(correction::compute_display_cols): New function.
(correction::ensure_terminated): Use new member name m_byte_length.
(line_corrections::add_hint): Add multibyte awareness.
(layout::print_trailing_fixits): Likewise.
(layout::get_x_bound_for_row): Likewise.
(test_one_liner_simple_caret_utf8): New self-test analogous to the one
with _utf8 suffix removed, testing multibyte awareness.
(test_one_liner_caret_and_range_utf8): Likewise.
(test_one_liner_multiple_carets_and_ranges_utf8): Likewise.
(test_one_liner_fixit_insert_before_utf8): Likewise.
(test_one_liner_fixit_insert_after_utf8): Likewise.
(test_one_liner_fixit_remove_utf8): Likewise.
(test_one_liner_fixit_replace_utf8): Likewise.
(test_one_liner_fixit_replace_non_equal_range_utf8): Likewise.
(test_one_liner_fixit_replace_equal_secondary_range_utf8): Likewise.
(test_one_liner_fixit_validation_adhoc_locations_utf8): Likewise.
(test_one_liner_many_fixits_1_utf8): Likewise.
(test_one_liner_many_fixits_2_utf8): Likewise.
(test_one_liner_labels_utf8): Likewise.
(test_diagnostic_show_locus_one_liner_utf8): Likewise.
(test_overlapped_fixit_printing_utf8): Likewise.
(test_overlapped_fixit_printing): Adapt for changes to
get_affected_columns, get_printed_columns and class corrections.
(test_overlapped_fixit_printing_2): Likewise.
(test_linenum_sep): New constant.
(test_left_margin): Likewise.
(test_offset_impl): Helper function for new test.
(test_layout_x_offset_display_utf8): New test.
(diagnostic_show_locus_c_tests): Call new tests.
gcc/testsuite/ChangeLog:
2019-12-09 Lewis Hyatt <lhyatt@gmail.com>
PR preprocessor/49973
* gcc.dg/plugin/diagnostic_plugin_test_show_locus.c
(test_show_locus): Tweak so that expected output is the same as
before the diagnostic-show-locus.c changes.
* gcc.dg/cpp/pr66415-1.c: Likewise.
From-SVN: r279137
Diffstat (limited to 'contrib/unicode/README')
-rw-r--r-- | contrib/unicode/README | 44 |
1 files changed, 44 insertions, 0 deletions
diff --git a/contrib/unicode/README b/contrib/unicode/README new file mode 100644 index 0000000..01ae2c1 --- /dev/null +++ b/contrib/unicode/README @@ -0,0 +1,44 @@ +This directory contains a mechanism for GCC to have its own internal +implementation of wcwidth functionality. (cpp_wcwidth () in libcpp/charset.c). + +The idea is to produce the necessary lookup table +(../../libcpp/generated_cpp_wcwidth.h) in a reproducible way, starting from the +following files that are distributed by the Unicode Consortium: + +ftp://ftp.unicode.org/Public/UNIDATA/UnicodeData.txt +ftp://ftp.unicode.org/Public/UNIDATA/EastAsianWidth.txt +ftp://ftp.unicode.org/Public/UNIDATA/PropList.txt + +These three files have been added to source control in this directory; +please see unicode-license.txt for the relevant copyright information. + +In order to keep in sync with glibc's wcwidth as much as possible, it is +desirable for the logic that processes the Unicode data to be the same as +glibc's. To that end, we also put in this directory, in the from_glibc/ +directory, the glibc python code that implements their logic. This code was +copied verbatim from glibc, and it can be updated at any time from the glibc +source code repository. The files copied from that respository are: + +localedata/unicode-gen/unicode_utils.py +localedata/unicode-gen/utf8_gen.py + +And the most recent versions added to GCC are from glibc git commit: +2a764c6ee848dfe92cb2921ed3b14085f15d9e79 + +Finally, the script gen_wcwidth.py found here contains the GCC-specific code to +map glibc's output to the lookup tables we require. This script should not need +to change, unless there are structural changes to the Unicode data files or to +the glibc code. + +The procedure to update GCC's wcwidth tables is the following: + +1. Update the three Unicode data files from the above URLs. + +2. Update the two glibc files in from_glibc/ from glibc's git. Update + the commit number above in this README. + +3. Run ./gen_wcwidth.py X.Y > ../../libcpp/generated_cpp_wcwidth.h + (where X.Y is the version of the Unicode standard corresponding to the + Unicode data files being used, most recently, 12.1). + +After that, GCC's wcwidth will match the most recent glibc. |