Arrow Research search

Author name cluster

Junfeng Pan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
1 author row

Possible papers

2

IS Journal 2007 Journal Article

Cost-Sensitive-Data Preprocessing for Mining Customer Relationship Management Databases

  • Junfeng Pan
  • Qiang Yang
  • Yiming Yang
  • Lei Li
  • Frances Li
  • George Li

A staged-framework for data preprocessing has been developed to support data mining and help service providers identify customers who might switch to a competitor. The framework pushes the cost sensitivity and data imbalance of customer retention data into the data preprocessing itself. Tests using data set from the ACM KDD Cup 1998 showed that the framework outperformed the winner of that data mining and knowledge discovery competition. The framework has also been incorporated into a software system, called ED-Money. To demonstrate the framework's ability to predict customer attrition with high accuracy, it was applied to some benchmark data and to a real customer attrition data set from a large Chinese mobile telecommunications company

AAAI Conference 2005 Conference Paper

Competence Driven Case-Base Mining

  • Rong Pan
  • Junfeng Pan

We present a novel algorithm for extracting a high-quality case base from raw data while preserving and sometimes improving the competence of case-based reasoning. We extend the framework of Smyth and Keane’s case-deletion policy with two additional features. First, we build a case base using a statistical distribution that is mined from the input data so that the case-base competence can be preserved or even increased for future problems. Second, we introduce a nonlinear transformation of the data set so that the case-base sizes can be further reduced while ensuring that the competence be preserved and even increased. We show that Smyth and Keane’s deletion-based algorithm is sensitive to noisy cases, and that our solution solves this problem more satisfactorily. We show the theoretical foundation and empirical evaluation on several data sets.

v2026.09.13