Rethinking Outliers: Ali Hadi on Fuzzy Logic, Data Precision, and Cluster Analysis

Rethinking Outliers: Ali Hadi on Fuzzy Logic, Data Precision, and Cluster Analysis

Dr. Ali S. Hadi, an outstanding Egyptian scholar in statistics, is a former distinguished visiting professor at the American University in Cairo and an emeritus professor at Cornell University.
He has published numerous articles in his areas of expertise: advanced statistics, data analysis, and applied statistics. He has specifically conducted research on outliers data points that lie outside the majority of the data in a particular data set.
 

MSTF Media reports:
Science and Technology Exchange Programs (STEP), held as part of Mustafa(pbuh) Prize Week, facilitate the circulation of scientific knowledge among Muslim scientists. The 10th STEP, thus, hosted a variety of international figures in science and technology with significant achievements. Among these figures was Ali S. Hadi, a former professor at the American University of Cairo.  
In an interview with the MSTF, he answered questions about his research over the years. Asked about the reason why relying on a crisp cutoff for outlier detection cannot always be reliable, Hadi stated that when collecting and analyzing data, outliers present one of the biggest challenges since they are extreme values that usually do not belong to the majority of the data.
“For example, when you are collecting data regarding people’s income, and suddenly Bill Gates’ data comes up, in this case, Gates would be considered an outlier in income because of the considerable chasm there is between his data and the others,” Hadi explained.
Hadi’s research focuses on multivariate data, which involves working with not just one variable for example, income but multiple variables. This, according to him, complicates the process of detecting outliers because you will not be able to graph them. Data can be graphed in one, two, or three dimensions, but anything above that would make things complicated. In this case, the statistics professor added, one has to rely on numeric techniques to identify outliers. 
“Some of the outliers will be just on the borderline between the normal and non-normal. You might declare them as outliers, you might not,” he maintained. “So, if you use a crisp cutoff point, you will make mistakes. I'm considering the values that are close to those crisp cutoff points.”
This is why statisticians use Fuzzy Logic to determine the degree of the outliers, which stands anywhere between zero and one. If it's closer to one, it will be declared an outlier. If it's closer to zero, it will be declared a non-outlier.
Hadi then addressed the benefits of using Fuzzy methods for outlier detection compared to the classic BACON algorithm. The original BACON algorithm, he stated, is based on crisp. But Hadi and his team have extended it to the Fuzzy BACON, which does not group data into outlier and non-outlier; rather, data at the borderline are evaluated in a fuzzy area. 
“The crisp model is a special case of fuzzy. So if you put the fuzziness parameter to zero or one, it becomes crisp. If you put the fuzziness anything between zero and one, it becomes fuzzy, and you can analyze data at the borderline with more precision.”
Regarding the difference between Fuzzy BACON and the classic version, Hadi thus elaborated: “In the classic model, there is a fixed cutoff point. Anything above is an outlier, and anything below, not an outlier. The fuzzy model, however, provides you with a range of values. Anything in this range is fuzzy in the sense that you cannot be certain whether it is an outlier or not. Certainty about whether it is an outlier or not is measured gradually.”
Pointing out the application of this model in various types of data, Hadi explained that the original model was designed for purely numeric data, not for variables that are categorical or nominal. There are new techniques that allow BACON to work on data of different types, including numeric as well as categorical data. The range of its applications has thus been extended.
The distinguished professor further mentioned one of his latest projects in an area called clustering. One of the fundamental questions in this study is how the distance between clusters can be measured with precision clusters are data collected based on similarity.
“So far, the literature on this area has proposed several methods,” Hadi explained. “The latest proposal is the Elliptical Distance.” 
While the classical Euclidean Distance assumes that variables have the same variance and are independent of each other, in reality, variables have different variances, and most of the variables are dependent on each other. The Elliptical Distance takes these facts into account. 
Elaborating on the advantages of his proposed method, Hadi pointed out two of its most significant features. “My proposed method recognizes two things: first, the variables’ different measurements and variance; second, their dependence on one another,” he maintained. 
In traditional models, variables with a big variance overwhelm those with a small variance. These models also do not take the dependence into account. Hadi’s proposed method overcomes the two disadvantages, increasing the precision of cluster analysis to a considerable extent.