[llvm] [CGData] Declare command line options in TableGen, one struct per library (PR #226087)

Fangrui Song via llvm-commits llvm-commits at lists.llvm.org
Sat Sep 26 10:47:45 PDT 2026


================
@@ -0,0 +1,51 @@
+//===----------------------------------------------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+include "llvm/Option/OptParser.td"
+
+def CGDataOptions : OptionsStruct;
+
+defm CodeGenDataGenerate : BoolField<"codegen-data-generate", "false",
+  "Emit CodeGen Data into custom sections">;
+defm CodeGenDataThinLTOTwoRounds : BoolField<"codegen-data-thinlto-two-rounds",
+  "false", "Enable two-round ThinLTO code generation. The first round emits "
+  "codegen data, while the second round uses the emitted codegen data for "
+  "further optimizations.">;
+defm CodeGenDataUsePath : ValueField<"codegen-data-use-path", "StringRef",
+  "", "File path to where .cgdata file is read">;
+defm GlobalMergingCallOverhead : ValueField<"global-merging-call-overhead",
+  "double", "1.0", "The overhead cost associated with each function call when "
+  "merging functions.">;
+defm GlobalMergingExtraThreshold : ValueField<"global-merging-extra-threshold",
+  "double", "0.0", "An additional cost threshold that must be exceeded for "
+  "merging to be considered beneficial.">;
+defm GlobalMergingInstOverhead : ValueField<"global-merging-inst-overhead",
+  "double", "1.2", "The overhead cost associated with each instruction when "
+  "lowering to machine instruction.">;
+defm GlobalMergingMaxParams : ValueField<"global-merging-max-params",
+  "unsigned", "std::numeric_limits<unsigned>::max()",
+  "The maximum number of parameters allowed when merging functions.">;
+defm GlobalMergingMinInstrs : ValueField<"global-merging-min-instrs",
+  "unsigned", "1",
+  "The minimum instruction count required when merging functions.">;
+defm GlobalMergingMinMerges : ValueField<"global-merging-min-merges",
+  "unsigned", "2", "Minimum number of similar functions with the same hash "
+  "required for merging.">;
+defm GlobalMergingParamOverhead : ValueField<"global-merging-param-overhead",
+  "double", "2.0", "The overhead cost associated with each parameter when "
+  "merging functions.">;
+defm GlobalMergingSkipNoParams : BoolField<"global-merging-skip-no-params",
+  "true", "Skip merging functions with no parameters.">;
+defm IndexedCodeGenDataLazyLoading : BoolField<
+  "indexed-codegen-data-lazy-loading", "false", "Lazily load indexed "
+  "CodeGenData. Enable to save memory and time for final consumption of the "
+  "indexed CodeGenData in production.">;
+defm IndexedCodeGenDataReadFunctionMapNames : BoolField<
+  "indexed-codegen-data-read-function-map-names", "true", "Read function map "
----------------
MaskRay wrote:

I'd switched to snake_case names and derived them from dashed names.

---

For snake_case:

- The mapping is trivial and lossless. `-` becomes _` `and nothing else changes. PascalCase has to guess at case, which is why codegen-data-thinlto-two-rounds needed an explicit CodegenDataThinLTOTwoRounds. With snake_case there's no case to get wrong, so the named-defm override would only be needed for spellings that can't be identifiers: a . in the name, a leading digit, or a C++ keyword.
- It lines up with OptTable. The member is x_y_z and the IDs are OPT_x_y_z and OPT_x_y_z_EQ, which is what OptTable users already read. The emitter would derive both from one function instead of the two it has now.
- It's greppable from the flag. A reader who sees -codegen-data-generate in a RUN line finds codegen_data_generate with a trivial substitution. Today they have to know it became CodegenDataGenerate, or CodeGenDataGenerate before this change.
- It matches how other compilers treat declared options. GCC exposes flag_foo_bar, and rustc exposes sess.opts.foo_bar. Both read like the flag.

Against:

- LLVM style. The coding standards want UpperCamelCase variables and members, and readability-identifier-naming will flag the reads. The fair counter: these are generated members named by the spelling, much as the OPT_ enumerators already break the enumerator rule. Expect a reviewer to raise it anyway, so the PR description should say so up front.
- MLIR uses camelBack members (disableThreading), so MLIR code would read Global.mlir_disable_threading. It's consistent within the framework but foreign to MLIR style.


https://github.com/llvm/llvm-project/pull/226087


More information about the llvm-commits mailing list