arXiv Science⌕ Search

arXiv subjects

Rohan Yashraj Gupta

Publications and source records attributed to Rohan Yashraj Gupta.

3 recordsLinked to original sources

TNF based Spectral Embedding for Effective Application of Supervised Machine Learning Techniques in Automobile Insurance Fraud Detection

Fraud detection is an important area of research in the insurance business due to its financial implications. The primary aim of a fraud detection model is to identify fraud and non-fraud cases with high accuracy along with other important metrics such as Sensitivity, Specificity, Precision, F1-score, False Positive Rate, False Discovery Rate, AUC etc. To achieve this, we need to explore a suitable classification model to identify fraud and non-fraud cases. In this work, we have used auto insurance data set and explored classification models such as Decision Tree (DT), Random Forest (RF), XGBoost, LightGBM and Gradient Boosting Machine (GBM). To overcome the problem of data imbalance, we have employed MWMOTE and TGAN techniques. We have used Topological Node Feature(TNF) based spectral embedding for low dimensional data representation along with some popular embedding methods like MDS, Isomaps and t-SNE. After studying all the 65 possible combinations of these models, we have proposed an innovative method for effective automobile insurance fraud detection. For the given dataset, our results show that using a combination of MWMOTE as a data imbalance handling technique (Phase I), TNFSE2 as data embedding (Phase II) and Random Forest as classification (Phase III) provides the best result in comparison to all other combinations. This work also highlights the efficacy of TNF based spectral embedding in automobile insurance dataset

cs.LG↗

Markov model with machine learning integration for fraud detection in health insurance

Fraud has led to a huge addition of expenses in health insurance sector in India. The work is aimed to provide methods applied to health insurance fraud detection. The work presents two approaches - a markov model and an improved markov model using gradient boosting method in health insurance claims. The dataset 382,587 claims of which 38,082 claims are fraudulent. The markov based model gave the accuracy of 94.07% with F1-score at 0.6683. However, the improved markov model performed much better in comparison with the accuracy of 97.10% and F1-score of 0.8546. It was observed that the improved markov model gave much lower false positives compared to markov model.

cs.LG↗

Implementation of Correlation and Regression Models for Health Insurance Fraud in Covid-19 Environment using Actuarial and Data Science Techniques

Fraud acts as a major deterrent to a companys growth if uncontrolled. It challenges the fundamental value of Trust in the Insurance business. COVID-19 brought additional challenges of increased potential fraud to health insurance business. This work describes implementation of existing and enhanced fraud detection methods in the pre-COVID-19 and COVID-19 environments. For this purpose, we have developed an innovative enhanced fraud detection framework using actuarial and data science techniques. Triggers specific to COVID-19 are identified in addition to the existing triggers. We have also explored the relationship between insurance fraud and COVID-19. To determine this we calculated Pearson correlation coefficient and fitted logarithmic regression model between fraud in health insurance and COVID-19 cases. This work uses two datasets: health insurance dataset and Kaggle dataset on COVID-19 cases for the same select geographical location in India. Our experimental results shows Pearson correlation coefficient of 0.86, which implies that the month on month rate of fraudulent cases is highly correlated with month on month rate of COVID-19 cases. The logarithmic regression performed on the data gave the r-squared value of 0.91 which indicates that the model is a good fit. This work aims to provide much needed tools and techniques for health insurance business to counter the fraud.

stat.AP↗