Hashing Modulo Context-Sensitive 饾浖-Equivalence
Abstract
The notion of 伪-equivalence between 位-terms is commonly used to identify terms that are considered equal. However, due to the primitive treatment of free variables, this notion falls short when comparing subterms occurring within a larger context. Depending on the usage of the Barendregt convention (choosing different variable names for all involved binders), it will equate either too few or too many subterms. We introduce a formal notion of context-sensitive 伪-equivalence, where two open terms can be compared within a context that resolves their free variables. We show that this equivalence coincides exactly with the notion of bisimulation equivalence. Furthermore, we present an efficient O(nlogn) runtime hashing scheme that identifies 位-terms modulo context-sensitive 伪-equivalence, generalizing over traditional bisimulation partitioning algorithms and improving upon a previously established O(nlog2 n) bound for a hashing modulo ordinary 伪-equivalence by Maziarz et al. Hashing 位-terms is useful in many applications that require common subterm elimination and structure sharing. We hav employed the algorithm to obtain a large-scale, densely packed, interconnected graph of mathematical knowledge from the Coq proof assistant for machine learning purposes.