Arrow Research search

Author name cluster

James Geller

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
1 author row

Possible papers

8

AIIM Journal 2015 Journal Article

A comparative analysis of the density of the SNOMED CT conceptual content for semantic harmonization

  • Zhe He
  • James Geller
  • Yan Chen

Objectives Medical terminologies vary in the amount of concept information (the “density”) represented, even in the same sub-domains. This causes problems in terminology mapping, semantic harmonization and terminology integration. Moreover, complex clinical scenarios need to be encoded by a medical terminology with comprehensive content. SNOMED Clinical Terms (SNOMED CT), a leading clinical terminology, was reported to lack concepts and synonyms, problems that cannot be fully alleviated by using post-coordination. Therefore, a scalable solution is needed to enrich the conceptual content of SNOMED CT. We are developing a structure-based, algorithmic method to identify potential concepts for enriching the conceptual content of SNOMED CT and to support semantic harmonization of SNOMED CT with selected other Unified Medical Language System (UMLS) terminologies. Methods We first identified a subset of English terminologies in the UMLS that have ‘PAR’ relationship labeled with ‘IS_A’ and over 10% overlap with one or more of the 19 hierarchies of SNOMED CT. We call these “reference terminologies” and we note that our use of this name is different from the standard use. Next, we defined a set of topological patterns across pairs of terminologies, with SNOMED CT being one terminology in each pair and the other being one of the reference terminologies. We then explored how often these topological patterns appear between SNOMED CT and each reference terminology, and how to interpret them. Results Four viable reference terminologies were identified. Large density differences between terminologies were found. Expected interpretations of these differences were indeed observed, as follows. A random sample of 299 instances of special topological patterns (“2: 3 and 3: 2 trapezoids”) showed that 39. 1% and 59. 5% of analyzed concepts in SNOMED CT and in a reference terminology, respectively, were deemed to be alternative classifications of the same conceptual content. In 30. 5% and 17. 6% of the cases, it was found that intermediate concepts could be imported into SNOMED CT or into the reference terminology, respectively, to enhance their conceptual content, if approved by a human curator. Other cases included synonymy and errors in one of the terminologies. Conclusion These results show that structure-based algorithmic methods can be used to identify potential concepts to enrich SNOMED CT and the four reference terminologies. The comparative analysis has the future potential of supporting terminology authoring by suggesting new content to improve content coverage and semantic harmonization between terminologies.

AIIM Journal 2009 Journal Article

Using WordNet synonym substitution to enhance UMLS source integration

  • Kuo-Chuan Huang
  • James Geller
  • Michael Halper
  • Yehoshua Perl
  • Junchuan Xu

Objective Synonym-substitution algorithms have been developed for the purpose of matching source vocabulary terms with existing Unified Medical Language System (UMLS) terms during the integration process. A drawback is the possible explosion in the number of newly generated (potential) synonyms, which can tax computational and expert review resources. Experiments are run using a synonym-substitution approach based on WordNet to see how constraining two methodological parameters, namely, “maximum number of substitutions per term” and “maximum term length, ” affects performance. Our hypothesis is that these values can be constrained rather tightly—thus greatly speeding up the methodology—without a marked decline in the additional matches produced. Furthermore, we investigate whether a limitation on only the first of the two parameters is sufficient to achieve the same results. Methods A four-stage synonym-substitution methodology using WordNet is presented. A group of experiments is carried out in which the two methodological parameters “maximum number of substitutions per term” and “maximum term length” are varied. The purpose is to examine their effect on the growth in the number of potential synonyms generated and the associated loss of results. The experiments are based on the re-integration of the “Minimal Standard Terminology” (MST) into the UMLS. Synonym-substitution matches found to be inconsistent with the current content of the UMLS and thus deemed to be incorrect are further manually scrutinized as an audit of the original integration of the MST. Results An increase of 11% in the number of “MST term/UMLS term” matches was achieved using the synonym-substitution methodology. Importantly, this result prevailed when tight threshold values (such as a maximum of two synonym substitutions per term) were imposed on the parameters. Furthermore, it was found that limiting only the “maximum number of substitutions per term” parameter was sufficient to obtain the performance enhancement. During the additional audit phase, a number of the reported mismatches were actually seen to be correct, representing an additional 10% increase in the number of matches obtained. Conclusion A synonym-substitution methodology that utilizes WordNet is a useful automated aide in UMLS source integration. Experiments showed that there was a significant speed-up but no degradation in match results when the methodology's “maximum number of substitutions per term” parameter was relatively tightly constrained. The methodology also helped to discover errors in the MST's original integration, and improve the quality of the UMLS's conceptual content.

AIIM Journal 2005 Journal Article

A lexical metaschema for the UMLS semantic network

  • Li Zhang
  • Yehoshua Perl
  • Michael Halper
  • James Geller
  • George Hripcsak

Objective: A metaschema is a high-level abstraction network of the UMLS’s semantic network (SN) obtained from a partition of the SN’s collection of semantic types. Every metaschema has nodes, called meta-semantic types, each of which denotes a group of semantic types constituting a subject area of the SN. A new kind of metaschema, called the lexical metaschema, is derived from a lexical partition of the SN. The lexical metaschema is compared to previously derived metaschemas, e. g. , the cohesive metaschema. Design: A new lexical partitioning methodology is presented based on identical word-usage among the names of semantic types and the definitions of their respective children. The lexical metaschema is derived from the application of the methodology. We compare the constituent meta-semantic types and their underlying semantic-type groups with the previously derived cohesive metaschema. A similar comparison of the lexical partition and a published partition of the SN is also carried out. Results: The lexical partition of the SN has 21 semantic-type groups, each of which represents a subject area. The lexical metaschema thus has 21 meta-semantic types, 19 meta-child-of hierarchical relationships, and 86 meta-relationships. Our comparison shows that 15 out of the 21 meta-semantic types in the lexical metaschema also appear in the cohesive metaschema, and 80 semantic types are covered by identical meta-semantic types or refinements between the two metaschemas. The comparison between the lexical partition and the semantic partition shows that they have very low similarity. Conclusion: The algorithmically derived lexical metaschema serves as an abstraction of the SN and provides views representing different subject areas. It compares favorably with the cohesive metaschema derived via the SN’s relationship configuration.

AIIM Journal 2005 Journal Article

An expert study evaluating the UMLS lexical metaschema

  • Li Zhang
  • George Hripcsak
  • Yehoshua Perl
  • Michael Halper
  • James Geller

Objective: A metaschema is an abstraction network of the UMLS's semantic network (SN) obtained from a connected partition of its collection of semantic types. A lexical metaschema was previously derived based on a lexical partition which partitioned the SN into semantic-type groups using identical word-usage among the names of semantic types and the definitions of their respective children. In this paper, a statistical analysis methodology is presented to evaluate the lexical metaschema based on a study involving a group of established UMLS experts. Methods: In the study, each expert was asked to identify subject areas of the SN based on his or her understanding of the various semantic types. For this purpose, the expert scans the SN hierarchy top-down, identifying semantic types, which are important and different enough from their parent semantic types, as roots of their groups. From the response of each expert, an “expert metaschema” is constructed. The different experts’ metaschemas can vary widely. So, additional metaschemas are obtained from aggregations of the experts’ responses. Of special interest is the consensus metaschema which represents an aggregation of a simple majority of the experts’ responses. Statistical analysis comparing the lexical metaschema with the experts’ metaschemas and the consensus metaschema is presented. Results: The analysis results shows that 17 out of the 21 meta-semantic types in the lexical metaschema also appear in the consensus metaschema (about 81%). There are 107 semantic types (about 79%) covered by identical meta-semantic types and refinements. The results show the high similarity between the two metaschemas. Furthermore, the statistical analysis shows that the lexical metaschema did not grossly underperform compared to the experts. Conclusion: Our study shows that the lexical metaschema provides a good approximation for a partition of meaningful subject areas in the SN, when compared to the consensus metaschema capturing the aggregation of a simple majority of the human experts’ opinions.

AIIM Journal 1999 Journal Article

A methodology for partitioning a vocabulary hierarchy into trees

  • Huanying (Helen) Gu
  • Yehoshua Perl
  • James Geller
  • Michael Halper
  • Mansnimar Singh

Controlled medical vocabularies are useful in application areas such as medical information systems and decision-support systems. However, such vocabularies are large and complex, and working with them can be daunting. It is important to provide a means for orienting vocabulary designers and users to the vocabulary's contents. We describe a methodology for partitioning a vocabulary based on an IS-A hierarchy into small meaningful pieces. The methodology uses our disciplined modeling framework to refine the IS-A hierarchy according to prescribed rules in a process carried out by a user in conjunction with the computer. The partitioning of the hierarchy implies a partitioning of the vocabulary. We demonstrate the methodology with respect to a complex sample of the MED, an existing medical vocabulary.

IJCAI Conference 1997 Conference Paper

Challenge: How IJCAI 1999 can Prove the Value of AI by Using AI

  • James Geller

The past several years have seen much progress in the area of propositional reasoning and satisfiability testing. There is a growing consensus by researchers on the key technical challenges that need to be addressed in order to maintain this momentum. This paper outlines concrete technical challenges in the core areas of systematic search, stochastic search, problem encodings, and criteria for evaluating progress in this area.

IJCAI Conference 1985 Conference Paper

The Teachable Letter Recognizer

  • James Geller

The "Teachable Letter Recognizer" (TLR) is a program for letter learning which uses a method not vet described in the character recognition literature. TLR has two main characteristics: (1) It uses a discrimination tree as a knowledge representation. The discrimination tree is a quadtree with some additions. (2) The two types of global features that are used to characterise a letter are the density of pixels and the overall "color", white, black or grey. TLR is invariant to shifts and shows several interesting effects which are related to human behavior, for example occasionally it becomes confused when learning new letters. TLR's importance as a letter recognition program lies in its ability to recognize some distortions of letters which it has never seen before, and for which it also does not have a transformation algorithm.

v2026.09.13