Flang improve fidelity of unformatted I/O endianness with FORT_CONVERT_UNIT (#223831)

Add further fidelity to specifying desired endianness of unformatted I/O
file on a per unit basis with the `FORT_CONVERT_UNIT` environment
variable.

To reviewers (friends) I'm not sure about:

1. The correct function prefixing for the FLANG runtime environment.
2. How to test/validate on Windows
3. How to test/validate with CUDA runtime support enabled

Thank you.


https://discourse.llvm.org/t/runtime-unformatted-i-o-conversion-between-big-and-little-endian-formats/91751

Documentation for new environment variable:

## `FORT_CONVERT_UNIT`

Determines the data conversion applied to specific unformatted I/O
units.

```
FORT_CONVERT_UNIT= mode | mode ';' exception | exception ;
mode: 'NATIVE' | 'LITTLE_ENDIAN' | 'BIG_ENDIAN' | 'SWAP';
exception: mode ':' unit_list | unit_list ;
unit_list: unit_spec | unit_list ',' unit_spec ;
unit_spec: integer | integer '-' integer ;
integer: [0-9]+ ;
```

The endianness of unformatted files is determined in the following
order:
1. The host processor's native endianness.
5. The setting of the `-fconvert=<mode>` flang compiler command-line
option.
6. The global setting from the `FORT_CONVERT=mode` or
`FORT_CONVERT_UNIT=mode`
environment variables. Contradictory mode settings between
`FORT_CONVERT` and
`FORT_CONVERT_UNIT` result in `FORT_CONVERT_UNIT` taking priority.
7. The explicit setting of the `CONVERT=` specifier in the `OPEN`
statement for
a particular unit.
8. The exception setting for an individual unit or range of units from
the
`FORT_CONVERT_UNIT` environment variable.

### Examples

* If the `FORT_CONVERT`, `FORT_CONVERT_UNIT`, and `CONVERT=` specifier
in the
  `OPEN` statement are all missing, no data conversion is performed for
  unformatted I/O; the host processor's native encoding is used.
* If the environment variable `FORT_CONVERT=BIG_ENDIAN` is set and no
`CONVERT=` specifier is present in the `OPEN` statement, input is
assumed to
  be big-endian, and output is emitted in big-endian format.
* On a little-endian host, if the environment variable
`FORT_CONVERT=SWAP` is
  set and the `CONVERT=LITTLE_ENDIAN` specifier is present in the `OPEN`
statement, the `OPEN` statement takes precedence: input is assumed to be
  little-endian, and output is emitted in little-endian format.
* If unit 10 is opened on a little-endian host with the environment
variable
  `FORT_CONVERT=SWAP`, `FORT_CONVERT_UNIT=BIG_ENDIAN:10`, and an `OPEN`
  statement for unit 10 with the `CONVERT=LITTLE_ENDIAN` specifier, the
`BIG_ENDIAN:10` exception from the `FORT_CONVERT_UNIT` environment
variable
takes precedence: input is assumed to be big-endian, and output is
emitted
  in big-endian format.

### Notes

1. `<mode>` values specified with the runtime environment variables
   `FORT_CONVERT` or `FORT_CONVERT_UNIT` are case-insensitive.
2. `<mode>` values specified with the flang command-line option
   `-fconvert=<mode>` are case-sensitive and include:
`mode: 'native' | 'little-endian' | 'big-endian' | 'swap';`
3. `unit_spec` supports ranges separated by a hyphen. Ranges must denote
positive unit numbers, and the starting unit (LHS) must be less than or
equal
to the ending unit (RHS).
4. Unit numbers and ranges can be specified multiple times with
different
`modes`, with the last (rightmost) `exception` taking priority. For
example:
`FORT_CONVERT_UNIT="LITTLE_ENDIAN:10,11,15-20;BIG_ENDIAN:19"`
The conversion for unit 19 will be `BIG_ENDIAN`.
5. If an `exception` is not prefixed with `mode:`, `mode` is assumed to
be
`BIG_ENDIAN`.  For example:
`FORT_CONVERT_UNIT="20-25;LITTLE_ENDIAN:26-27"`
Regardless of the host processor's endianness, units 20 through 25 will
be
treated as big-endian for both input and output, while units 26 and 27
will be
treated as little-endian for both input and output.

---------

Co-authored-by: Slava Zakharin <szakharin@nvidia.com>
7 files changed
tree: 2b0f49ea4b3f06231033ab5212774a99386e0f3f
  1. .ci/
  2. .github/
  3. bolt/
  4. clang/
  5. clang-tools-extra/
  6. cmake/
  7. compiler-rt/
  8. cross-project-tests/
  9. flang/
  10. flang-rt/
  11. libc/
  12. libclc/
  13. libcxx/
  14. libcxxabi/
  15. libsycl/
  16. libunwind/
  17. lld/
  18. lldb/
  19. llvm/
  20. llvm-libgcc/
  21. mlir/
  22. offload/
  23. openmp/
  24. orc-rt/
  25. polly/
  26. runtimes/
  27. third-party/
  28. utils/
  29. .clang-format
  30. .clang-format-ignore
  31. .clang-tidy
  32. .git-blame-ignore-revs
  33. .gitattributes
  34. .gitignore
  35. .mailmap
  36. CODE_OF_CONDUCT.md
  37. CONTRIBUTING.md
  38. LICENSE.TXT
  39. pyproject.toml
  40. README.md
  41. SECURITY.md
README.md

The LLVM Compiler Infrastructure

OpenSSF Scorecard OpenSSF Best Practices libc++

Welcome to the LLVM project!

This repository contains the source code for LLVM, a toolkit for the construction of highly optimized compilers, optimizers, and run-time environments.

The LLVM project has multiple components. The core of the project is itself called “LLVM”. This contains all of the tools, libraries, and header files needed to process intermediate representations and convert them into object files. Tools include an assembler, disassembler, bitcode analyzer, and bitcode optimizer.

C-like languages use the Clang frontend. This component compiles C, C++, Objective-C, and Objective-C++ code into LLVM bitcode -- and from there into object files, using LLVM.

Other components include: the libc++ C++ standard library, the LLD linker, and more.

Getting the Source Code and Building LLVM

Consult the Getting Started with LLVM page for information on building and running LLVM.

For information on how to contribute to the LLVM project, please take a look at the Contributing to LLVM guide.

Getting in touch

Join the LLVM Discourse forums, Discord chat, LLVM Office Hours or Regular sync-ups.

The LLVM project has adopted a code of conduct for participants to all modes of communication within the project.