Class for storing and accessing the IR2Vec vocabulary.
The Vocabulary class manages seed embeddings for LLVM IR entities. The seed embeddings are the initial learned representations of the entities of LLVM IR. The IR2Vec representation for a given IR is derived from these seed embeddings.
The vocabulary contains the seed embeddings for three types of entities: instruction opcodes, types, and operands. Types are grouped/canonicalized for better learning (e.g., all float variants map to FloatTy). The vocabulary abstracts away the canonicalization effectively, the exposed APIs handle all the known LLVM IR opcodes, types and operands.
This class helps populate the seed embeddings in an internal vector-based ADT. It provides logic to map every IR entity to a specific slot index or position in this vector, enabling O(1) embedding lookup while avoiding unnecessary computations involving string based lookups while generating the embeddings.