CIS 660

Data Mining with Advanced Machine Learning (4-0-4)

Course Content

  • Class Announcement and Post
  • Class Syllabus
  • Lab Assignments
  • Projects
  • Class Lecture Notes


  • Class Announcement and POST





    Semester Schedule:
    See University's Official Academic Calendar for the Semester Schedule to Add and Drop and the Final Exam Schedules



    19. August 25, 2025:

    The CIS660 Final Exam Info:

    Final Exam schedule for Fall 2025:
    Wed Dec 10, 12:30PM - 2:30PM in Class


    For the Active Lecture Notes Links for the CIS660 Final Only, Access the following site: The rest of links are all disabled.

    the Active Lecture Notes Links Only for the CIS660 Final Only

    One Page Note is Allowed to the Final.

    The Exam Format and Rules will be the same as the Midterms (Short Exams).

    Subjects to Focus On for the Final

    Important Problems to Resolve to Increase Model Accuracy of Classification
    Clustering Algorithms - K-Mean, Heirarchical Clustering Algorithms, DBScan,
    Accuracy Estimation and Validation, Model Evaluation Methods, Model Comparison Methods (NOT for this Semester)
    Advanced NLP: Word2Vec, BERT
    RNN, LSTM
    Deep Learning Architecture with CNN
    (Not for This Semeser):
    Association Rule Mining: Apriori Algorithm and Optimization, Assocaition Rule Generation, Frequent Pattern Tree, Interesting Measures




    18. August 25, 2025:

    Midterm Info:

    Midterm Info and The Lecture Note Links Only for Midterm

    There Will Be 2 Exams (15% Each)

  • Exam 1
  • Exam 1 will be Tentatively on the Last Week of September !

    Subjects to Focus on: 

    Chapter 2, 3 on :
    Basic Stats of Data, All the Data Proximity Measures, Data Transformation Methods, Feature Selection Methods,
    Feature Correlation Measures -- Kai Square, Correlation, Covariance
    Basic Text Analysis Algorithms and Methods in Information Retrieval: TF-IDF Measure, Document as Term Vector


    Note that You are responsible ONLY for the Lecture Notes and the Subjects Covered in Class

    Question Types:
    6-7 Short Answer Questions on Problem Solving with Given Small Data Sets
    No T/F Questions, No Multiple Choice






  • Exam 2
  • Exam 2 will be Tentatively on the Last Week of October !

    Subjects to Focus on: 

    The Second Exam on Oct 30 on Machine Learning Algorithms.
    The Subjects Covered on Machine Learning Algorithms and Their Objective Functions Covered in Class by Mon Oct 28:
    All the Lectures on Decision Tree, K-NN, Naive Bayes, Baysian Belief Network, Ensemble Methods
    Linear Regression, Logistic Regression, Multinomial Logistic Regression, ANN Feed Forward, Backpropagation Algorithms,
    Ensemble Methods -- Bagging, Boost, Ada Boost, Sampling Techniques, Model Evaluation Methods 
    Common Problems in Classification and ML Models, Data Preprocessing Methods of Each Classifier


    The Same Question Types as the Exam 1




    15. August 25, 2025:

    There Will be Random Quizzes to Check the Attendance !
    Only 5 - 10 Mins Will Be Allowed for Each Quiz. Those Who Come to the Class Late, They Will NOT Be Given a Quiz !  


    13. August 25, 2025:

    Only the registered students can access the course blackboard.
    If you have a problem with your blackboard access, please contact the registrar or CSU Tech Support to resolve the issue !

    Center-for-elearning Technical Support

    Faculty does not control your registration and the course blackboard access in the CSU systems.

    3. August 25, 2025:

    TA Information:


    TA:
  • Jiacheng Jerry Guo

  • Email: j.guo58@vikes.csuohio.edu

    Office Hours: Mon and Wed 4:00 pm -6:00 pm (Send him email ahead to let him know you are coming)

    Location: FH311 or ZOOM Meeting br>
    ZOOM Meeting ID:
    Passcode:








    If you have questions in Labs or Grading your Labs, Send an email to TA or Talk to him During his TA Office Hours or Schedule a Zoom meeting with him






    Dr. Chung's Office Hours:

    Tues and Thursday 1:30PM - 3:30PM
    Office: FH 222 or Zoom Meeting
    Email: s.chung@csuohio.edu
    Send Me Email to Set Up a Meeting or Zoom Meeting.

    Meeting ID: 859 3867 6332
    Zoom Meeting Link




    2. August 25, 2025:

    The Output of each lab is your report in Doc file that shows your screen captures of your data processing steps for analytics and the results
    Each of your screen capture must show each step and the result returned by the server in the SAME window in your System to prove that your lab is done correctly in YOUR SYSTEM !!

    Lab Submission:
    1. Submit your Lab in zip file including 1) your lab report in .doc and 2) all the source files, preprocessed input files, outputs on Blackboard for a timestamp as a proof.

    2. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Heading of the Front Page of Your Report in Bigger and Bold Font !




    1. August 25, 2025:
    The class webpage link is announced in the Class Blackboard as well.

    Class Webpage


    Semester Schedule:
    See University's Official Academic Calendar for the Semester and the Final Exam Schedules

    Final Exam schedule for Fall 2025:
    Wed Dec 10, 12:30PM - 2:30PM in Class


    Midterm Info:

    Midterm Info and The Lecture Note Links Only for Midterm

    There Will Be 2 Exams (15% Each)

  • Exam 1
  • Exam 1 will be Tentatively on the Last Week of September !

    Subjects to Focus on: 

    Chapter 2, 3 on :
    Basic Stats of Data, All the Data Proximity Measures, Data Transformation Methods, Feature Selection Methods,
    Feature Correlation Measures -- Kai Square, Correlation, Covariance
    Basic Text Analysis Algorithms and Methods in Information Retrieval: TF-IDF Measure, Document as Term Vector


    Note that You are responsible ONLY for the Lecture Notes and the Subjects Covered in Class

    Question Types:
    6-7 Short Answer Questions on Problem Solving with Given Small Data Sets
    No T/F Questions, No Multiple Choice






  • Exam 2
  • Exam 2 will be Tentatively on the Last Week of October !

    Subjects to Focus on: 

    The Second Exam on Oct 30 on Machine Learning Algorithms.
    The Subjects Covered on Machine Learning Algorithms and Their Objective Functions Covered in Class by Mon Oct 28:
    All the Lectures on Decision Tree, K-NN, Naive Bayes, Baysian Belief Network, Ensemble Methods
    Linear Regression, Logistic Regression, Multinomial Logistic Regression, ANN Feed Forward, Backpropagation Algorithms,
    Model Evaluation Methods 
    Common Problems in Classification and ML Models, Data Preprocessing Methods of Each Classifier


    The Same Question Types as the Exam 1




    Important Note for Exams:
    I am afraid that it is not possible to change the schedule of the midterm or Final exam for one person's favor.
    I can NOT make anyone take the Midterm or Final of CIS660 individually (in Person or in Remote). There are many reasons due to serious security bleaches and cheating related concerns. The class size is too big, I simply cannot allow any exception for an individual reason.
    There is always a possibility that exams of 2-3 courses would be scheduled to a same day during the Midterm or Final Exam Weeks. The midterm date of CIS660 is usually decided considering many factors such as the progress of class subjects and the progress of the majority of the class students. It will be announced at least 2 weeks ahead.







    The prerequisite of CIS660 has been changed to CIS530 and CIS550, which was intended to prevent the students without any CS/DS Undergrad Background from Registering for CIS660 in their first semester without completing preparatory courses. 
    CIS530 and CIS550,Engineering Statistics are Required, and an Undergrad (Intro) of Machine Learning or Algorithms Are Preferred as Prerequistes of CIS660.

  • Class Syllabus of CIS 660 Data Mining with Advanced Machine Learning
  • Class Syllabus of CIS660/CIS760 Data Mining with Advanced Machine Learning

  • Lab Assignments



  • Labs:

    First Week Lab0: Choose your System/Tool/Platform to Set Up and Get Used to:

    See the Lab Assignment0 Section Below to See the Set Up Guide for Python Data Science Platform in Step by Step !


    Basic Python Tutorials :

  • Python Tutorial
  • Python Codecademy Tutorial Site


  • Python Data Science Platforms

    There Are Mainly Three Ways to Set up All the Neccessary Data Science Software/Library below !

    See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !!
    Set Up Guide for Python Data Science Platform in Step by Step ************************************


    Additional Installation Guide:
    Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system)


    Python IDE Deduggers:

    Basic Guide for Python Debugger Pycharm
    Python IDE Spyder
    Python Debugger Spyder



    Python Data Science Platforms:

    Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system)

    Basic Guide for Python tools for Data Science: jupyter-notebook
    More Basic Guide for Python tools for Data Science: jupyter-notebook


    • Anaconda
    Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms
    Anaconda Tutorials

    • Python Anaconda Tutorial Sites
    Anaconda Tutorial Site
    Anaconda Tutorial Site

    • PyTorch
    PyTorch Site (It can be integrated from Anaconda as well)

    Google Colab:
  • Google Colab


  • • Python Scikit Learn for Common Data Science Tasks
  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing
  • Python Text Processing Libs for Text Analysis
  • Python Numpy Tutorial

  • Text Preprocessing (Natural Language Processing) Library in Python SpaCy:

    Liquistic Modules in Python SpaCy
    Lemmatizer in Python SpaCy
    Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy
    How to Code Liquistic Modules like Lemmatizer in Python SpaCy
    Python Example for Basic Text Processing

    • Python sklearn.cluster Python Sklearn Clustering






    There is Another Data Science Platform in R if You Choose to Learn (SAS Data Miner or MS Data Tool Have an Integrated R Platform)

    Basic R Tutorials :

  • R Studio basic Tutorial
  • R Basic Online Lecture

  • R Manuals

  • R Tutorial with Examples with R Stat Tool
  • Note that Examples in this tutorials may not the final correct output for Lab1 !
  • Examples of Basic R Stat Tool with Helpful References

  • Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study by Nick White (Now in FaceBook)







    Useful Machine Learning Tutorial Sites:

    Keras for Image Processing/Text Processing with Deep Learning

    Google Colab for FAST Machine Learning Execution in GPU


    Good Data Science and AI Platforms for Developing and Training:
  • Open AI
  • Google Colab
  • Tutorial: Getting Started with Google Colab
  • Medium by MIT




  • For Your Own Study
  • Coursera Machine Learning by Stanford



  • Lab Submission Instructions:


    The Output of each lab is your report in Doc file that shows your screen captures of your system/tool/platform setting/configurations, your data processing steps for analytics and the results
    Each of your screen capture must show each step and the result returned by the server in the SAME window in your System to prove that your lab is done correctly in YOUR SYSTEM !!

    1. Submit (on Blackboard) your Lab in a zip file including 1) your lab report in .doc and 2) all the source/scipt files, preprocessed input files, outputs on Blackboard for a timestamp as a proof.
    2. If you did an Extra Credit Lab, Make a Note on the COVER of Your Lab Clearly !

    See the Instructions on How to Create Your Lab Report Below !


    The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard !
    You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission

    If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font !

    Always Follow the Deadline of Each Lab Assigned on the Class Blackboard.
    The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester.

    Please Identify Your Course When You Ask Me in Email !



  • The Output of each lab is your Lab Report in Doc file that shows your screen captures of each of your executions with Your Outputs

  • Your Report in Doc file should include all the platform set up procedures, the execution steps, and copy of each source code files
  • Each of your screen capture must show your results returned by your experiment to prove that you have done the lab correctly !!



  • 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof.

    2. IMPORTANT NOTE:
    Your Screen Captures in Your Lab Report Should Show Your Own System Name to Prove That Your Lab Was Done In Your System.

  • Instructions on How to Create Your Lab Report
  • How to Create Table Contents in a word Doc File for Your Lab Report
  • Example of Output of the Execution Steps for Labs

  • 3. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font !

    4. Always Follow the Deadline of Each Lab Assigned on the Class Blackboard.
    The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester.




  • The Output of each lab is your Lab Report in Doc file that shows your screen captures of each of your executions with Your Outputs

  • Your Report in Doc file should include all the platform set up procedures, the execution steps, and copy of each source code files
  • Each of your screen capture must show your results returned by your experiment to prove that you have done the lab correctly !!



  • 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof.

    2. IMPORTANT NOTE:
    Your Screen Captures in Your Lab Report Should Show Your Own System Name to Prove That Your Lab Was Done on Your Computer.

  • Example of Output of the Execution Steps for Labs

  • 3. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font !

    4. Always Follow the Deadline of Each Lab Assigned on the Class Blackboard.
    The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester.









    Lab Assignment 0:


    Set Up Your Python Data Science Platform for Labs and Final Projects by the End of the First Weekend !

    There Are Mainly Three Ways to Set up All the Neccessary Data Science Software/Library !

    See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !!
  • Set Up Guide for Python Data Science Platform for Three Options in Step by Step ************************************



  • Optional for Big data Processing
  • See Lab Section of CIS593 Big Data for More Detailed Guides





















  • Although the Database Skills and Knowledge are not required for this course, it will be very useful for handling big data sets.
    For Those Who want to Use a Database Server for Initial Data Handling

    See the Lab Sections of CIS 430/530 for SQL Server Installation Guides
    See the Lab Sections of CIS 611 for SQL Server/Data Warehouse OLAP Server Installation Guides










    REQUIRED LAB ASSIGNMENTS:


    The Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard !


    Lab Assingment 1:


  • Lab Assignment 1 on Data Preprocessing and Transformation, Similarity Measure and Correlation Matrix

  • Lab1_1 - Part 1 and Part 2 on Preprocessing and Transformation: Due By the End of the Third Week
    Lab1_2 - Part 3 and part 4 on Similarity Measure and Correlation Matrix: Due By the End of the Fourth Week


    You don't Have to Import AdventureWork Data Warehouse to Get the View vTargetMailCustomer for Lab1. You can directly download vTargetMailCustomer.csv below.

    Input Data File for Lab1:
  • vTargetMailCustomer Added Outliers Income & Age (csv) Customer Profile Data Set
  • Lab1 Small Data Set (Mouse Serum Data) for Null Replacement (csv) (Three Columns Marked for Null Replacement)


  • TA Instruction on How to Generate Lab1 Report


  • For Part 1 and 2:

    For each selected feature, identify ALL the required data preprocessing methods like Normalization, Discretization, Binarization, and more based on the feature data properties. They must be done along with other preprocessing methods.

    Examples of Lab1 Report of Data Preprocessing and Data Similarity Measures

    Example1 of Lab1 Output: Part 1
    Example of Lab1 Output: Similarity Measure

    Note that these Examples Here May NOT neccessarily all correct ! They just show How Lab1 can be done as an example. Do NOT blindly follow !

    For example, EnglishEducation should be Transformed as Categorical or Ordinal?
    Depending on the properties of the column, the corresponding transsformation method should be applied.




    For Part 3 and 4:

    IMPORTANT NOTES !!!
    You Are Not Supposed to Use any Python Libs to Calculate Similarity Measures to Build a Similarity Matrix or Correlation to Build a Correlation Matrix in Part 3 and Part 4 of Lab 1.
    You have to write Scripts/Programs to Compute each Similarity or Correlation and Build a Similarity Matrix or Correlation Matrix.
    If you use the built-in lib to build a Correlation Matrix, which is one line of code, you will get 0 for the part.




    Suggested Platforms to Use for Lab1:

  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing

  • Anaconda

    Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms
    Anaconda Tutorials


  • FAQs on Lab 1
  • Tutorial on Data Preprocessing: One Hot Encoder
  • Data Normalization and Binarization Encoding Example for Neural Net classifier
  • Final Selected Feature Set



  • Lab Assingment 2:

  • Lab Assignment 2 on Text Analysis
  • Note That Lab2 Requires Stemming/Lemmarizaion and Bi-Gram Handling

  • Common NLP Preprocessing Tasks to Be Done for TextAnalysis
  • FAQ on Lab2
  • Example Lab2 Output Structure(Note That This May NOT Show Correct Output Values of Lab2)

  • You can use any avaialble API such as Beautiful Soup for cleaning webpages in html tags

  • Set Up Beautiful Soup for Web Scraping for Text Analysis
  • Documentation of Beautiful Soup
  • How to Remove HTML Tags from Webpages for NLP Processing of Text Analysis

  • Some Trouble Shooting Tips:
    For Lab 2, it requires the python library 'clean-text'
    The full command to install this dependency is:
    pip install clean-text
    If you do wish to rerun it and make sure that it works, you would have to uninstall the old one first as they use the same module name.
    Having both installed at the same time will result in Python using the incorrect one.


    NOTE !!!
    You Are Not Supposed to Use any Python Libs to Build a Cosine Similarity Matrix for Lab 2.
    You have to write a script to compute and build a Similarity Matrix.
    If you use the built-in lib to build a Cosine Similarity Matrix, you will get 0 for Labs.




    Lab Assingment 3 on Classification

  • Lab Assignment 3 on Classification with Machine Learning

  • IMPORTANT NOTES for LAB3:

    1. Use the Best Selected Features Given Below instead of Your Own Selected Feature Set in Lab1 :
  • Target Mail Selected Feature Set

  • IMPORTANT !!: In your Lab report, Make sure to Show the final transformed training set file and all the attribute values for the first two objects
    in your training set that was used for your classifiers.


    2. Choose your Classifiers for Lab3 as below:

  • Two Classifiers from the Probabilistic Approach Based ML Algorithms
    1) One from Decision Tree or Bayesian and
    2) Another from Any Ensemble Methods: Random Forest, XGBoost, LightGBM, Adaboost

  • Two Classifiers from the Numerical Approach Based ML Algorithms
    3) ANN
    4) SVM
    or 4) One from Distance Based K-NN


  • 3. Grading is Based on Your Best Accuracy with the best input hyperparameters you identified from Your Experiments









    Lab Assingment 4 on Clustering:

  • Lab Assignment 4 on Clustering
  • 1. Choose 2 data sets given in this section and Do Clustering with K-Mean and DBScan Algorithms.
    2. Experiment to find the best Input Parameters for each Algorithm.
    3. For each Clustering result in your experiment, Apply any method discussed in the Lecture notes (Ward's method, Silhouette score, Elbow method, Entrophy/Purity) to Measure the quality of the Clustering result.
    4. Visualize the final best clusters
  • FAQs for Lab 4 on Clustering
  • Data Preprocessing Issues for Clustering
  • Tutorial on Advanced Clustering Analysis with Mixed Data Types



  • Data Sets you can Choose for Lab4:

  • DataSet: Customer Creditcard Usage GENERAL
  • You can do random sampling on NASA HTTP Server Log to create a smaller data set to apply clustering.
  • NASA Webserver Log file Description and Download site (NASA HTTP Access Logs) (Scroll Down to the bottom)
  • NASA HTTP Access Logs in compressed file


  • Some Suggested Platforms for Clustering

    scikit-learn Clustering


    Although Clustering in general should work on data as multidimensional vectors,
    For NIJ Data For Lab4, you can work on NIJ Challeneges to identify Hot Spots for Crime Location - GPS data and Crime Category instead of handling data as multidimensional object for Clustering algorithms
    Cluster on the Crime Loactions (X and Y coordinates - transform them to GPS coordinates) with K-Mean and DBCSAN and Cluster the Crime locations for each Crime Category
    but Add PAI or PEI Analysis for Hot Spot Analysis with Changing Parameters

    National Institute of Justice (NIJ) Crime HotSpot Analysis Challenging:


    NIJ Crime Location Data Sets:

  • NIJ Challenging Data Published in 2017
  • More Recent Data Sets in NIJ Challenging Data Site

  • NIJ (National Institute of Justice) Crime Forecasting Challenge: data description and overview
  • PAI: How to Measure HotSpot Cluster of NIJ GIS Data in PAI
  • Tutorial: How to Measure HotSpot Clusters of NIJ GIS Data in PAI

  • How to Visulaize NIJ GIS Data in QGIS
  • How to Visulaize NIJ GPS data with ArcGIS



  • AWID data Set: EXTRA CREDIT (see more info in the Clustering and Anomaly Detection Section)

    AWID Site:
    AWID: Intrusion Detection in Wireless Network Server Log Data

    Apply a Feature Selection Method Discussed in class to reduce the dimensions to apply clustering

  • AWID Data Set Create a Smaller Data Set Only for Lab4 for Clustering

  • Suggest to Use EmEditor to open a big file if WordPad can't open it.
  • How to Open and Handle the AWID Data File on Window

  • Sometimes a data file collected from a different file system (HDFS, for example) has incompatible special characters for line feed or some others to open it in a Window/Linux system, so you have to write and run a simple script to replace those special characters in the file.
    See a sql script to import the AWID to Sql server
  • Sql Script to Import the AWID file to SQL Server

  • IEEE Paper Description of AWID Data Set and Attack Types Make Your Own Data with Random Sampling on Normal Type and Select only Two Different Attack Types for Clustering



  • PROJECT



    Project:

  • Project and Paper List (To Be Updated) See More updated List Below
  • Project Resources and Reserach Paper List on GPT Training Methods with Context Information for Medical Question Answering
  • See CIS612 Big Data Project Site for Big Text Data Sets



    Important Dates for Final Group Project

  • Task 1: Group Project Proposal Due by Nov 9 !

    Submit Your Group Proposal (in Minimum 3 Pages) on BlackBoard

    Group Proposal Should Include:

    1. Data Description, Data Size, Data Collection Plan (if needed),
    2. Systems/Tools to Use,
    3. Data Preprocessing Methods,
    4. Data Analytic Goal with Plan of Evaluation (Design of Your Experiment) in Detail




  • Task 2: Group Project Status Report ! Due By Nov 21

    Your Project Status Report Should Show the Following Tasks Done:

    1. Platform Setting/Configuration Procedure if it is new,
    2. Your Data Contents, Selected Feature Description,
    3. Data Preprocessing Steps and the intermediate Outputs




  • Task 3: Project Presentation Starts on Last Week of the Semester Monday Dec 1 !
  • Read the Instructions of Project Presentation Here ! (This Google Project Presentation Scheduler Will Be Sent To Your CSU Email !)

    Group Project Presentation Should Include:

    1. Data Description, Data Size, Data Collection Method(if needed),
    2. Systems/Tools to Used,
    3. Feature Selection Method and The Final Features Selected
    4. Data Preprocessing Methods and Intermediate Results, Final Training Set and Test Set Description,
    5. Data Analytic Goal with (Design of Your Experiment) in Detail
    6. The Problems/Errors Encountered and Your Resolutions
    7. Evaluation Results and Visualization of the Results




    Any Group Size in 1 - 4 Person Group Are Allowed
    Note that Large Groups (3-4 Person Group) Shoud Complete a Bigger Project !


    One Submission Per Group Is Required !
    EACH Memeber Name and ID SHOULD BE LISTED in the COVER !
    List Your First and Last Name Only ! Exactly as Appeared in the CSU CampusNet. DO NOT Use Your Middle Name.

    If We Can't Find You by Your Name appeared on Your Project Report and Presentation Schedule. Your Project will be Considered as Missing with 0.

     



    Task 4:
    Final Group Project Report Submission Instructions (By the End of Friday of Your Presentation Week):

    Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week !

    Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx !

    If your data file is too big to upload, Submit your zip file with your Data file on your ONE drive and Send email to me and TA attached the data file on One Drive.




    Submit a Zip file that includes:

    1. All of your presentation slides (both in .ppt and .pdf) and
    2. Your Group Final Project Report (in doc) that shows with:

    1.Platform/System Set up Procedures/Instructions
    Make Sure to List the Specific Version of Each Component of Your Platform (Such As the Versions of Python, GPU, Pretrained Deep Learning ML Model, and more) to Run Your Project.

    2. Description of the Analytic Goal of Your Project and Description of the Data Set
    2. I Need to See Your Transformed Features for each ML Algorithms and Your Source Scripts/Codes.

    3. Evaluation Results and Visualization of the Results
    4. Executions Steps, all the source codes/scripts, all the intermediate outputs, and final output files
    5. Include the Problems/Error Encountered and Your Resolutions in Your Report

    All These Are Required Becasue I Need to See Proof/Evidence Showing that Your Project Is Not a Copy of a Github Codes You Downloaded from the Internet.



    One Submission Per Group Required.

    Submit a Zip File that Includes All the required Source Files, Input, Output Files, and Final Report (in Doc file) and Presentation Slides (in pptx).

    Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files.

    The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results.

    If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes from the Web.







    If you need high computing power system with GPU for training:

    1. Good Platforms for Training to Develop Deep Learning Models or Large Language Models: Free GPU Use
  • Google Colab
  • Tutorial: Getting Started with Google Colab
  • Medium by MIT
  • Open AI

  • 2. Big Data Servers with GPU are Available to Use in the Big Data Lab: (Temporary Permission will be given per request)

    3. High Computing System in the Engineering College:
  • Guide site for Request Permission to Use GPU-based High Performance Computing System at CSU
  • README






    Good Data Sets for Classification:

    Health Data Sets for Challenges ****************
    US Government Data Sets ****************



    More Project List and Data Sets:

    Sentiment Analysis of Online Reviews/Social Media data:
  • Trip Advisor Hotel Review Data Set (from Carnegie Mellon Data set)
  • Amazon Product Review Data Set (from UCSD Research Site) Request data sets on category: Mobil/Cell Phone Reviews Or Electronics
  • IMDB Moview Review data sets
  • NLP Text Data Set Repository


  • Research Direction/Guides for Project:

    Question Answering System/Smart/Intelligent System on Texts/Papers/Documents

  • Project Resources and Reserach Paper List on GPT Training Methods with Context Information for Medical Question Answering

  • Recent Research Papers on Large Language Model Training Methods :

    Prompting Based Training:
    Research on Prompting Based Training


    Recent Papers:
    BIOGPT
    ChatGPT 3.5 Natively Performs Chain of Thoughts 2023
    Least to Most Prompting 2023
    RankVicuna: Reranking GPT Prompt Model 2023

    Paper Topic Summarization Weak Supervision CMU 2020


    Data Set for Question Answering Training

    Pubmed Paper Repository Data Sets
    Pubmed QA Data Set at HuggingFace
    Pubmed QA

    NLP Data Sets






  • Good Conference Site List to Search Research Papers on Review Analysis (From UCSD Research Site)


  • How to Collect Social Media Data: (take CIS612 for More on Big Data Processing and NoSQL Big Data Systems)

    This can be done by collecting Twitter real time stream and apply the basic NLP and Text analysis methods for Sentiment Analysis.
    You have only about 10 days to collect the data. You have to apply the developer's account to teh Twitter to start.
    However, it requires some knowledge and skills to collect and process the big data stream in the Twitter logging strucrure
    before you start any text analytic tasks, which are not the focus of CIS660 because of the extremely limited time to cover all the important analytic subjects.
    I suggest you to take CIS612 for those subjects. Most of CIS Master/PhD students take both CIS612 and CIS660.


    CIS 612 (or CIS 593 Big Data) cover all the related subjects for the social media opinion analysis for:
    1) How to collect the real time big data like social media Twitter and
    2) how to process and manage the collected big data using semi structured database server for further data analytic process.
    You can try to start collecting the data. Look at the instructions and step by step tutorials in Lab3 section in CIS593 website below.

    Lab Section of CIS 593 Big data for Twitter Data Collection (Scroll down to the Lab3 Section)

    Tutorial: How to Collection Twitter Messages to Process with MongoDB and Tableau
    Tutorial: How to Get FaceBook Graph API Data to Analytics with Hive


    General Task Guideline for Sentiment Analysis of Social Media Texts

    General Steps for Sentiment Analysis


    Sample Projects on Sentiment Analysis

    Basic Sentiment Analysis with Yelp Review Data
    Basic Another Sentiment Analysis with Yelp Review Data
    ProjectExample on LDA for Topic Discovery and Word2Vec for Similar Words with Trip Advisor Hotel Review Data


    Lectures to Learn Methods and Find Related Research Papers:
    Paper Presentation on Mining and Summarizing Feature Reviews (KDD 2004)
    Research Summary on Product Feature Reviews by Dr. Bing Liu
    Research Summary on Sentiment Analysis (Standford)


    Sentiment Analysis Project based on Earlier Papers:
    Research Paper: Sentiment Analysis Using Classification
    Research Paper Presentation on Sentiment Analysis Using Classification (by Xiaodan Liu)
    Research Paper: Sentiment Analysis Using Subjectivity






    Build a Ranking system or Recommendaton System on Product Reviews or Using Real Estate Land Use Data Sets:

    Using the Real Estate Land Use Data Sets and Working People Demographics in the Cleveland area (Data sets will be given)
    Generate a Ranking or Recommendation List for a Given User Interest or Prefernece on Real Estate Land Use Data Sets

    See More Info and Research Papers of Recommendation Systems/Ranking system in the Lecture Notes Section on Recommendation Systems After the Classification Section


  • Amazon Product Review Data Set for Recommendation Systems (from UCSD Research Site)



  • Project for Recommendation System with Real Estate Land Use Data Set in Cuyahoga County

    Meta data description is gievn below and the detailed instructions for the access to Data sets will be given later.
  • Data Description of Real Estate Land Use Data Set (The deatil instruction will be given)

    Permission to access this folder below will be given per request if your group want to do a Project on this data

  • Real Estate Land Use Data Set in Cuyahoga County (CSV files) (The deatil instruction will be given)
  • Real Estate Land Use Data Set in Cuyahoga County (in bak files to import to SQL Server)
  • How to Import bak File to SQL Server
  • Real Estate Land(Property) Use Data Set in Cuyahoga County (GEOJSON file) (The deatil instruction will be given)
  • GEOJSON file Data Description of Real Estate Land(Property) Use Data Set in Cuyahoga County (The deatil instruction will be given)
  • Processed Land Characteristic Data Sets of Real Estate Land(Property) Use Data Set in Cuyahoga County
  • Raw Characteristics Data Set of Each Land(Parcel) in Cuyahoga County (~200MB Zip file of CSV files ) See the data description file first.

  • How to Import bak File to SQL Server
  • How to Import bak File Manually to SQL Server from Remote Access



    Implement Your Own (Deep) Neural Network Architecture Using Google Tensorflow or MS CNTK or any of your choice for Text Analysis Task or Image Processing: Extra Credit !

    Extra Credit for Those Who Took CIS612 Big Data and HDFS:
    Do Your NN Training in Parallel on Hadoop. Research and learn how to train NN in parallel on HDFS. How to Combine each local model to one final one.


    If you want to work on Neural Network, then download source codes of skipgram (the first paper) from one of those sites below to learn how to build Skipgarm with NN then implement paragraph vector as document vector in the second paper.
    Look at the first research project for guide for this project at:
  • Project: Implemnting Paragraph Vector using Word2Vec

    Then, let's read the next two papers to understand the source codes for you to download and do the experiment with real data set.
  • Word2Vec Research Paper from Google
  • Paragraph Vector for Sentences and Document from Google 2014

    Then check these two sites where you can download all the source codes to start an experiment with real data sets.
  • Word2Vec to download (google site)
  • Word2Vec to download with npmjs
    Data set to train to generate word2Vec and Paragraph Vector:
    You can choose any webpage set or papers as data but it should be at least 200,000 documents to train
    There is a preprocessed wiki page data set is available in the Stanford NLP site that can be used as your training data set. You can make this as your project if you want.


    The Most Recent and Most Superior Word Vector: BERT -- See Advanced Text Analysis Section in Class Lecture Notes for more on BERT

    BERT from Google AI

    Tutorials on BERT from Google


    Good Tutorial Sites to Learn How to Use Pretrained BERT:

    SpaCy Embeddings for BERT Transformer

    Good Tutorial Site to Use Pretrained BERT Transformer

    Good Tutorial to Use Pretrained BERT Transformer for Question Answering System

    Good Tutorial to Use Pretrained BERT Transformer for Domain Specific System

    Good Tutorial Site to Use BERT Transformer for Sentence Classification



    Research Paper: Passage Reranking Using BERT from Google


    Anomaly Detection with AWID Data set:

    Implement one of methods in the recent papers on Anomaly Detection.

    See More Help on AWID data processing in Lab4 Section !

    AWID Site:
    AWID: Intrusion Detection in Wireless Network Server Log Data
    AWID Data Set Avaliable here: AWID: Wireless Network Server Log Data (1 GB zip)

    Paper that describe the data set:
    Paper: draft Intrusion Detection in 802-11 Networks Empirical Evaluation of Threats

  • Related Research Papers on Anomaly Detection


    NSL-KDD data set for AnomalyDetection:

    Kaggle: KDD data set for AnomalyDetection
    Data set Related Research site on AnomalyDetection


    For NIJ Challenging: Hot spot Analysis: See Clustering Section for more detailed information

    See More Info for NIJ Challegeing in the Clustering Section

    Download New data sets and Submission files of Winning Teams below to Learn from the NIJ (National Institute of Justice) Challenege site (Scroll down to see the links)
    Find Out How to Measure Hot Spots in score type PAI and PEI* for Every Category, Crime Type, time Frame.
  • NIJ Challenging with Hot Spot Analysis
  • NIJ Challenging Data from 2017
  • Tutorial: How to Measure HotSpot Cluster of NIJ GIS Data in PAI
  • How to Measure Cluster of NIJ GIS Data
  • Documenetation in Depth: How to Approach for Geospatial Data Analysis
  • Examples of Solutions: How to Use Challenges to Find Solutions

  • GIS Data Visualization API:
  • How to Visulaize NIJ GIS Data in QGISby Asanka Mananayaka
  • NIJ GPS data with ArcGIS


  • List of Top Journals and Conferences in Data Mining and Information Processing

  • R Tutorial for Project
  • Project Guide with an Example: Data Mining over LinkedIn Data
  • Project Guide with an Example on Twitter Text Mining
  • Presentation Schedule and List of What to Present in Your Presentation
  • How to read and Present a Research Paper


  • Some of Research Projects Done in Big Data Analytics Lab:

    Research Project: Document Clustering Using Word2Vec and Paragraph Vector

    Research Project: Machine Learning Based Full Text Intelligent Search Engine

    Research Project: Document Search Engine

    Research Project: Sentiment Analysis on Social Network Twitter

    Research Project: Feature Extraction and Selection Using AWID Data


    Project Presentations Examples

    Sentiment Analysis on Yelp Review Data

    Project: Text Mining for Sentiment Analysis: Predicting Review Stars (1 - 5) from Yelp Review Text data
    Project: Yelp Challenge with Text Mining: Predicting Review Stars (1 - 5) from Yelp Review Text data
    How to Get FaceBook Graph API Data to Analytics with Hive

    Research Papers Based on:
    Research Paper: Sentiment Analysis Using Classification
    Research Paper Presentation on Sentiment Analysis Using Classification (by Xiaodan Liu)
    Research Paper: Sentiment Analysis Using Subjectivity


    Text Analysis Using Classification
    Text Mining on Twitter Data on ISIS terrorists group and the fan groups of ISIS Santosh Tankala, Vishnu Vishnuteja Thummanapelli, and Akhi Reddy Laxmanagari
    Tutorial: How To get Twitter Data
    Paper Based On:
    Research Paper: Topic sentiment analysis in twitter: a graph-based hashtag sentiment classification approach


    Another Projects:

    Sentiment Analysis of Online Reviews:
  • IMDB Moview Review data sets
  • Trip Advisor Hotel Review Data Set (from Carnegie Mellon Data set)
  • Amazon Product Review Data Set (from UCSD Research Site)


  • 2. Recommendation System on Amazon Product Data set

    Project: Collaborative Filtering Algorithm Implementation for for Recommendation System Suhua Wei
    Project: Recommendation System Using Amazon Customer Product Data Set for Collaborative Filtering Sagar Dahiwala
    Research Papers Based on:
    Research Paper: Collaborative Filtering for Amazon Recommendation System /a>
    Paper: Recommendation System from Stanford Project Report
    Paper PPresentation: Amazon Recommendation System for Collaborative Filtering

    3. AnomalyDetection Using Classification:
    Intrusion Detection AnomalyDetection on Network Server Data by Sean Riehl and Yanan Lyu

    4. Question Answering Systems:
    < href="QASystemImageTextSagarMyur.pdf"> Project: Question Answering System on Image and Text Data Mayur and Sagar

    5. Clustering
    Independent Study Project: Outlier Detection on AWID Data Set Ahmad Arida
    Project: Outlier Detection on NASA HTTP Logs Ahmad Arida
    Project: Clustering and Hot Spot Analysis on Portland Map with NIJ Data Sarvesh Chande

    6. Data Analytics on Social Media Data
    Project: Data Mining over LinkedIn Data Danielle Aring and Noreen Halley
    Research Paper: A Tool for Collecting Provenance Data in Social Media







    Research Paper Presentation Guide (Extra Credit)

    For the extra credit Research paper presentation, you have to read and summarize the main method of the paper from the best conference sites in the Project List.

      Submit the slides in pptx and the paper. If you don't summarize the main content of the paper properly, no grade will be given.

    How to Read and What are the Important Contents to Summarize:

    How to read and present a research paper

    Example of Research Paper Presentation Slides

    More Research Paper Presenation examples






    Textbook Search Site at CSU the Library Site

    Web Accessible CIS660/EEC525 Textbook at the CSU Library Site


  • Class Lecture Notes with Tentative Schedule

    Class Chapter / Topic / Specific Objectives / Activities
    1

    Introduction to Big Data Analytics:

    Lecture Notes_1_1: Simple Overview of Big Data Analytics

    Lecture Notes: Introduction to Big Data, Big Data Processing and Big Data Analytics

    Lecture Notes_1_3: Big Data, Data Scientist and Research on Big Data Analytics at CSU


    1-4


    What is Data Mining?

    Chapter 1: Overview of Data Mining: What is Data Mining and Data Mining Steps


    Basic Data Structures and Data Properties in Data Mining:

    Chapter 2_1: What is Data Structure and Properties of DATA in Data Mining? (Kumar's Chapter 2_1)

    Chapter 2_1: Your DATA in Properties and Structures for Data Mining (Part 1) (J Han's Chapter 2_1)



    Example of Data Preprocessing and Transformation Steps For a Data Mining with Classification


    Data Preprocessing Steps: 1. Feature Selection, 2. Preprocessing:Cleaning, 3. Exploring to Get to Know Data, 4. Reduction/Integration, 5.Transformation, 6. Correct Measuring


    1. Feature Selection Methods: (To Be Covered at the end)





    2. Data Cleaning:

    Chapter2_1: Data Quality and Data Cleaning (Kumar's Chapter2_1)

    Chapter 2_2 : Knowing Your Data with Basic Statistics - Part 1(J Han's Chapter2_2)

    Chapter3_1: Data Cleaning for Null and Outliers in Data Preprocessing (J Han's Chapter3_1)



    Tutorials for Data Cleaning:

    - Null Value Replacement Methods:

  • Replace with Mean/Median/Min/Weighted Mean
  • Repalce with the Most Probable Vaue Inferenced from Probability Distribustion of each Value in each Feature

    - Outlier Removal/Replacement Methods:

  • 1.5IQR Method
  • Z-Score Method; Histogram
  • Inference with Regression Method
  • Many Other Advanced Methods: MCD(Minimum Covariance Determinant), KMI (Kenel Least Square)

  • Tutorials with Examples:
    Simple Examples of 1.5IQR for Outlier Removal Method with Box Plot per each Column

    Common Outlier Removal Methods for Each Column: 1.5IQR, Z-score, Histogram, Percentile

    Common Outlier Removal Methods for Each Column: 1.5 IQR, Histogram, Z-Score, Inference with Regression

    1.5 IQR for Outlier Removal with Box Plot per each Column


    Advanced Outlier Detection/Removal Methods:

    Python Tutorial: Advanced Outlier Removal Methods with Multiple Columns together

    Research Paper: Minimum Covariance Determinant: Advanced Outlier Removal Methods with Multiple Columns together

    Research Paper: Kernel Least Square : Advanced Outlier Removal Methods with Multiple Columns together






    3. Getting to Know Your Data

    Getting to Know Your Data with Basic Stats

    Chapter 2_2 : Knowing Your Data with Basic Statistics - Part 1(J Han's Chapter2_2)

    Visualization of Data Distribution of Each Feature in Data Set:

  • Box Plot
  • Histogram
  • Z-Score

  • Grouped Data Calculation




    Chapter 3: Overview of Data Exploration (Kumar's Chapter 3)


  • Well-known Probability Density Distribution Functions

  • How to Identify the Data Distribution of a Feature(Column) Values

    Gaussian (Normal) Distribution Function
    Log-Normal Distribution Function
    Beta Distribution Function
    Gamma Distribution Function







    4. Data Integraton and Reduction :

    Chapter 2_3: Intro to Data Integration, Reduction Methods - Aggregation, Discretization, and Sampling Methods (Kumar's Chapter 2_3)

    Data Cleaning and Reduction with Discretization Methods: Binning, Histogram (From the slides of J Han's Chapter 3_3)




    - Sampling Methods for Data Reduction:

    Bootstrap Sampling with Replacement
    Stratified Sampling with Weighted Mean


    - Oversampling and Resampling Methods for a Small Data Set:

    Stratified Bootstrap Sampling

    Monte Carlo Sampling for Each Feature

    More Advanced Resampling Methods in the Classification Section !















    5. Data Transformation: Normalization, Discretization, Binarization

    Chapter 3_4: Data Transformation - Normalization, Discretization (J Han's Chapter 3_4)

    Chapter 2_2: Data Transformation Basics (Kumar Chap2_2)



    Common Data Transformation Methods for Data Measures:

  • Binarization (Hot Encoding) Method for Categorical (Nominal) Data
  • Discretization for Data Reduction for Numeric Data with Too Many Unique values
  • Normalization/Standardization for Numeric Data



  • Tutorials for Data Transformation Methods:


    Correct Transformation Methods for ML Algorithms:

    Numerical Approach Based ML Algorithms -- ANN or SVM Requires Data Preprocessing with :

  • Normalization
  • Binarization for Any Categorical Attributes


  •    Tutorial for Data Preprocessing (Transformation): Normalization and Binarization (One Hot Encoding) for ANN or SVM

    Binarization for Categorical Features :

    Tutorial: Three Ways to Transform a Categorical Feature with Binarization (One Hot Encoding)

    Examples of Transformation Methods for ANN (Artificial Neural Network) with Binarization (One Hot Encoding) for Categorical Data and Normalization



    How to Transform TimeStamps:

    Examples of Preprocessing DateTime as Features for Classification: Appointment Cancellation Prediction

    Transformation Methods for Date Time Data Type

    Transformation Methods for Delta Time from Date Time Data Type





    Practical Problems and Limitations with Binarization (One Hot Encoding) for Categorical Data Transformation ***************





    Probablistic Approach Based ML Algorithms -- Decision Tree or Random Forest Reqires Data Preprocessing with :

  • Discretization for Any Continuous or Integer Attributes with Too Many Distinct Values

  • Discretization Methods and Its Limitations For Classification: Equal Width or Equal Depth (Frequency)




     




     

    Lectures on Data Reduction and Dimentionality Reduction Methods:

    Dimensionality Reduction Methods for Feature Selection: (Will Be Covered More after 5. Transformation and 6. Measures below)

    Lecture Notes with Examples of Dimensionality Reduction

    Chapter 2_4: Data Reduction Methods and Feature Selection Methods (Kumar's Chap2_4)

    Chapter 3_2: Data Reduction, Feature Selection Methods, and Sampling Methods (J Han's Chapter3_4)










    6. Data Measures:

    Proximity Measures:

    Chapter 2_4: Data Measures (Kumar's Chapter 2_2)

    Chapter_2_5: Knowing Your Data (Part2): Measuring Data Proximity Per Data Properties (J Han's Chapter2_3)

    Tutorial to Undersatnd Mahalanobis Distance



    Correlation Measure:

    Chapter 3_1: Correlation Analysis for Feature Reduction and Feature Correlation (J Han's Chapter3_3)


    Lecture on Covariance Matrix

    Tutorial for Covariance Matrix








    Other Supplementary Lecture Notes on Measures

    Advanced Mutual Info Mesures


    Scores and Ranking















    Dimensionality Reduction Methods for Feature Selection

    Lecture Notes on Curse of Dimentionality and Overview of Dimensionality Reduction Methods



    Supplementary Lectures on PCA (Principal Component Analysis)

    Tutorial on PCA (Principle Component Analysis)

    Lecture on PCA (Principle Component Analysis) Factor Analysis

    Lecture on Covariance and PCA (Principle Component Analysis) Analysis

    Wiki Tutorial for Covariance Matrix

    Singular Values of a Matrix




    Tutorials on PCA in Python/R:

    Example of PCA in Python with Data Set

    Tutorial for PCA Transformation in R

    Tutorial for Single Value Decomposition in R

    Full (YouTube) Lecture on Dimension Reduction Methods and PCA (Principal Component Analysis)






    Advanced Feature Selection Methods:

    Recursive Feature Elimination (RFE)


    Tutorials for RFE


    scikit-learn Feature Selection Methods


    Research on Advanced Feature Selection Method:

    Research in Big Data Lab: Examples of Feature Extraction and Feature Selection Methods with Dimensionality Reduction techniques


    4-5



    Unstructured Text Processing for Text Analytics

    Lexicon Based Approach with Document(Text) as Bag of Words Model:

    Lecture Notes on IR TF-IDF for Text/Web Mining (From Stanford Lecture Note) ******************

    BM 25 Ranking Algorithm (Parametrized TF-IDF with Document Length Normalization) ******************
    Lecture Notes_18: Lecture Notes on Basic Information Retrieval: Texts/Web Mining with Positioning Index for Phrase Query

    Lecture Notes_18: Lecture Notes on Inverted Index



    Google NGram Viewer : Inverted Index for all N Gram words -- Earlier Solution to N Grame (Phrase) Words

    Example Project of Building Inverted Index for Association Rule Mining in MS Analysis Service




    Letcures on Text Analysis Applications:

    Overview of Question Answering Systems (AI)

    Lecture Notes: Question Answering System and Query Analyzer

    Lecture Notes: Building IBM Watson IBM Watson System (AI):


    Sentiment Analysis:

    Lecture Notes: Overview of Sentiment Analysis System

  • Example Project on Twitter Text Mining and Sentiment Analysis

  • Sample Projects on Sentiment Analysis with Lexicon Based Bag of Words Model

    Basic Sentiment Analysis with Yelp Review Data
    Basic Another Sentiment Analysis with Yelp Review Data


    Common Natural Language Processing (NLP) Techniques for Text Preprocessing Tasks:

  • Common NLP Preprocessing Tasks to Be Done for Bag of Words Model with TF-IDF Scorng Algorithm for Text Analysis

  • Note that You Have to Decide Which Text Cleaning Preprocessing Should Be Applied or Should Not Be Applied in Your Text Processing Data Pipeline Depending on Your Goal of Text Analysis

    For TF-IDF Based/Lexicon Based Text Vectorization, You Can Apply All the Text Cleaning Preprocessing listed above.
    On the other hand, for the Rest of the NLP Parsers(API) such as POS Tagger, NER Tagger, You Should Not Remove Each Sentence Strcuture, Should Not Apply Some of the Common NLP Cleaning Tasks Such as Stop Word Removal Which Will Destroy a Sentence.

    Note that Removing Stop Words is Only for Lexicon (TF-IDF) Based Text Analysis. Removing Stop Words Could Have a Similar Effect (but Not Equivalent to) as Inverted Document Frequency(IDF) when You Cannot Measure IDF.
    However, It May NOT Be Good For Context Aware (Sentence Level) Information Extraction Methods like POS Tagger and NER Parser, or Word Embedding based Word2Vec Model Where Each Sentence Needs To Be Preserved

    See Annotator Dependencies Section in Standford Core NLP Site Below to Learn Required Preprocessing to Build a Correct Sequence of Your Data Pipeline ! *****


    NLP Preprocessing: Stemming-Lemmatization for Text Analysis:

  • Common NLP Preprocessing Task: Stemming-Lemmatization from Stanford NLP Group
  • Common NLP Preprocessing Task: Stemming-Lemmatization in python

  • See Python spaCy Libs for stemming and lemmatization below !



  • General (Short) NLP English Word List for Lexicon Based Text Analysis Applications:

    Stop word List
    Positive word List
    Negative word List


    To Get Lexicon Word (Dictionary) Lists
    WordNet -- Largest Lexical Database of English
    WordNet and Synset Download





    Summary of Text Similarity Measuring Methods:

    Text Similarity Measuring Methods










    Problems (Limitations) with Lexicon Based Approach with TF-IDF Ranking Score for Text Analysis:

  • Phrase (N-Gram Word) Identification
  • Can NOT Identify Relationships among Terms - Synonyms (Similar Meaning) or Antonyms (Opposite Meaning) of Terms
  • Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence - Polysemy Problem !
  • The order of Terms(Words) in a Sentence Is NOT Considered !!
  • Can NOT Identify New Terms or Changining Relationship Between New Terms and Old Terms


  • Solutions for the Probelm that TF-IDF Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence

    Context Aware Natural Language Processing (NLP) Methods:

    POS (Part of Speech) Tagging:

    Lecture Notes on POS (Part of Speech) Tagging (Stanford)
    English Grammar
    UPenn Treebank POS Tagging List




    NLP NER (Named Entity) Tagging:

    Lecture Notes on NER (Named Entity Recognizer) (Stanford)


    Software NER Tagging Parser:

    Stanford NLP NER (Named Entity) Tagging


    Application: Research in AI in Biomedical Domain

    Stanza NER Tagger on Medical text documents

    STANZA NLP NER (Named Entity) Tagging on Biomedical Documents (Pulished in 2021)





    CoreNLP Run Test the Examples of NLP Parsers Here !



    Important Text Preprocessing Data Pipeline for Text Analysis:

    See Annotator Dependencies Section to Learn Required Preprocessing to Build a Correct Sequence of Your Data Pipeline ! *****

  • Data Pipeline to Build for Text Preprocessing (by Stanford Core NLP APIs)
  • Stanford Core NLP APIs for Preprocessing/Parsers/Taggers to Build Text Processing Data Pipeline





  • a Good Research Paper on Training Medical Documents with Knowledge Graph (KDD 2019)





    Other Useful Core NLP Preprocessing APIs, Parsers (Taggers/Annotators) to Download to Use

  • Core Stanford Parser Download for POS and NER tagging (Github repo site as well)
  • Core NLP Preprocessing/Parsers/Taggers
  • NLTK (Natural Language Tool Kit)



  • SpaCy for Text Processing NLP Libs in Python :

    You can Also Use Python NLP Lib called SpaCy to build a Data Pipelining with NLP APIs for Text Analysis

  • Download and Installing spaCy
  • Introduction to spaCy101 for Text Proceesing with NLP Methods
  • spaCy: Building Data Pipleines for Text Proceesing with NLP Methods
  • Python spaCy Objects for Text Processing in Data Pipelines
  • Tutorials of spaCy for Text Proceesing with NLP Methods



  • Stanford NLP POS (part Of Speech) Tagging
    Stanford NLP Parser
    Stanford NLP POS Tagger in JAVA





    Syntactic Dependency Parser:

    Lecture Notes on Dependency (Stanford)
    Dependency Parser to Download (Stanford)


    Software Open IE (Information Extraction) Parser:

    Stanford NLP Open IE (Information Extraction) for Triple Extarction


    WordNet Site (by Princeton) to downlaod NLP Software:

    WordNet by Princeton
    Download Synnet/WordNet for Sentiment Analysis on Aspect/Feature Opinion Analysis


    For Those Who Want to Do a Research Project on Text Mining and Sentiment Analysis:

    General Guideline for Sentiment Analysis Tasks

    Fundmental Papers/Lecture Notes to Learn Basic Methods:

    Presentation: Summarizing Features of Product Reviews for Sentiment Analysis
    Paper: Summarizing Features of Product Reviews for Sentiment Analysis on : Example Project

    Pang's 2002 Paper: Sentiment Analysis on IMDB : Example Project
    Presentation: Sentiment Analysis on IMDB (From Pang's 2002 Papers)


    Research Summary on Aspect Opinion Analysis:

    Summary of Research Areas on Sentiment Analysis on Aspect/Feature Opinion Analysis
    Summary 2 of Research Areas on Sentiment Analysis on Aspect/Feature Opinion Analysis


    LDA Algorithms for Aspect/Feature Extraction for Topic Modeling:

    Document Topic Discovery Algorithm: Latent Dirichlet Allocation(LDA)
    LDA Paper:

    Presentation of LDA Based Topic Modeling Papers
    LDA Paper for Topic Modeling

    Basic Text Analysis Related Papers to Learn:

    Probabilistic Latent Semantic Indexing (Berkeley)
    LDA (Berkeley)


    Useful Tutorial Site for Download LDA:
    LDA Gensim for Topic Discovery
    LDA Gensim for Topic Discovery
    LDA gensim Download to try for Unsupervised Document Topic Identification




    Summary of Text Similarity Measuring Methods:

    Text Similarity Measuring Methods






    Semi-Supervised Learning Methods for Sentiment Analysis

    Well Known Ranking Algorithms for Vectorization of Each Document for Ranking or Labeling Ground Truth of a Training Set for a Machine Learning Classifier):

    Sample Lexicon Based Sentiment Analysis Proejct:
    2020 Presidential Election Prediction from Social Network Twitter Analysis





    Sentiment Analysis:

    Lecture Notes: Overview of Sentiment Analysis System


    Sentiment Analysis 2019

    General Guideline for Sentiment Analysis Tasks






    For Semi-Supervised Learning

    Well Known Ranking Algorithms for Sentiment Analysis

    Senti-WordNet for General English Texts from Google Page Ranking Algorithm

    Senti-WordNet for Sentiment Analysis
    Presentation of Paper: Senti-WordNet for Sentiment Analysis
    Paper: Senti-WordNet for Sentiment Analysis





    Page Rank

    Wiki site on Page Rank Algorithm
    Tutorial: Page Rank Algorithm
    Wiki site on Page Rank Algorithm




    VADER: Twitter Text Specific Ranking Algorithm

    Ranking Algorithm: VADER for Sentiment Analysis
    Compound Score of VADER Ranking Algorithm for Sentiment Analysis
    Paper: VADER for Sentiment Analysis




    Natural Language Processing (NLP) with Machine Learning Lectures Will Be Covered (See Below) After Advanced Classifications Are Covered !





    6-10



    Machine Learning for Data Anaytics

    Supervised Learning

    Classification with Machine Learning:


    Complete List of scikit-learn Supervised Learning Algorithms




  • Probabilistic Approach Based Machine Learning

  • Well-Known Optimized Decision Tree Based Machine Learning Algorithms

  • Random Forest (Based on CART and Bagging)
  • XGBoost (Extreme Gradient Boost)
  • Light GBM (Gradient Boosting Machine)
  • Ensemble Method: Adaboost


  • Lectures:

  • Decision Tree:
  • Lecture Notes_10: Lecture Notes on Basic Classification and Decision Tree by Kumar Book

    Corrections in Lecture Notes !

    SplitINFO as Complexity Penalty

    Lecture Notes_12: Lecture Notes on Basic Classification by J. Han Book

    Lecture Notes_10_1: Lecture Notes on CART wth Case Study

    Example of Decision Tree for Fault Detection with Overfitting Problem and Post Pruning


    scikit-learn Python Implementation of Decision Tree:
    scikit-learn: Classification with Decision Tree




  • Naive Bayes, Ensemble Methods, and Distance Based KNN

  • Lecture Notes: Alternative Classification: KNN, Naive Bayes, Ensemble Methods (Kumar Book)

    Lecture Notes: Advanced Classification: Baysian Belief Network, Regressions, (J. Han Book)

    Example of Bayesian Belief Networks
    Example of Bayesian Belief Networks





  • Distance Based KNN

  • K-Nearest Neighbor Algorithm

    Lecture on Nearest Neighbor Method



    scikit-learn site on Naive Bayes:
    scikit-learn: Classification with Naive Bayes

    Research Paper on Effect of Incorrect Categorical Data Encoding for Naive Bayes Classifier



  • Ensemble Methods:

  • Lectures:

    Lecture Notes on Ensemble Methods with Bagging and Boosting, AdaBoost, Random Forest





    Ensemble learning Methods:

    Ensemble Algorithm: Random Forest vs Bagging

    XGBoost: Exereme Gradient Boost Ensemble Algorithm


    scikit-learn Python Implementation of Ensemble Methods:

    scikit-learn: Ensemble Algorithm List

    Ensemble Method: Random Forest (based on Bagging with CART)

    scikit-learn: Advanced Ensemble Methods

    scikit-learn: Other Randomized Tree Ensemble Algorithms







    Ensemble Methods for Multi-Class Classification:

    One-vs-Rest (OVR), One-vs-One (OVO):

    One-vs-the-Rest or One-vs-One on Multi-Class Classifications
    Wiki site on One-vs-the-Rest(OvR) or One-vs-One(OvO)


    scikit-learn One-vs-Rest and One-vs-One

    scikit-learn: One-vs-Rest (OVR) Ensemble Methods
    scikit-learn: One-vs-One (OvO) Ensemble Methods








    Machine Learning with Numerical Approach:

  • Regression
  • Artificial Neural NetWork
  • SVM (Support Vector Machine)






  • Review on Regression for a Prediction Model:

    Lecture Note: Linear Regression
    Lecture Note: Linear Regression (Standford Tutorial)
    Introduction of Regression Models




    Standard Notations for Deep Learning


    Lectures:


  • Logistic Regression and Classification

  • Lecture on Logistic Regression

    Lecture on Logistic Regression and Gradient

    Lecture on Logistic Regression (Standford Tutorial)

    Lecture on Multinomial Logistic Regression (Standford Tutorial)




    Supplementary:

    Tutorial of Cross Entropy

    Softmax Function Tutorial

    Softmax Function Definition for Multinomial Classification




  • Artificial Neural NetWork


  • Lecture Notes on Neural Network to Know Better (Standford Tutorial)

    Tutorial with Example: Backpropagation in MultiLayer Neural Networks


    Matrix Implementation: ***********************************************

    Lecture: More Detail on Neural Network with Matrix Implementation of Backpropagation Algorithm (From DeepLearing.AI by Stanford Coursera)


    Practical Problems of Gradient Descent:

    Lecture on Gradient Descent Search and Regularization

    Gradient Descent
    Batch Gradient Descent vs Stochastic Gradient Descent



    scikit-learn: Classification with Neural Network

    How to Train NN When You Don't Have a Large Training Data Set:
    Difference between iterations with batch size and epochs



    Tutorial and Sample Codes for Implemenation of BackPropagation of NN and Classification Using Matrix Operations in Python

    Introduction to Deep Neural Network Architectures (To Be Covered in Deep Learning Section Below)















  • Support Vector Machine (SVM)


  • Lecture Notes on Concept of Artificial Neural Network, Support Vector Machine (SVM) (Kumar's Textbook)

    Lecture Notes on Concept of Artificial Neural Network, Support Vector Machine (SVM) (J Han"s Textbook)

    SVM Theory: Linear Classification
    SVM Theory: Non-Linear Handling
    SVM Theory: Non-Linear Classification with Kernel Functions



    Tutorial with Simple Examples:
    SVM Preprocessing Guides and Kernel Functions
    SVM Solutions When Not Linearly Separable

    scikit-learn SVM

    scikit-learn: Classification with SVM

    scikit-learn: SVM SVC

    SVM Kernel Functions

    Kernel Functions for SVM












    Tutorials:

    Tutorials for Data Preprocessing Methods:



    Practical Problems and Limitations with Binarization (One Hot Encoding) for Categorical Data Transformation ***************


    Discretization Methods and Its Limitations For Classification: Equal Width or Equal Depth (Frequency)




    ML Algorithms -- ANN or SVM Requires Data Preprocessing:
    Normalization for Any Numerical Attributes
    Binarization (One Hot Encoding) for Any Categorical Attributes

       Step by Step Implementation with Basic Examples for Lectures

    Walk Through with Basic Example using Panda for Classification with Neural Network

    Understanding SVM Algorithm Deeper for Classification with Python Example
    Walk Through with Basic Example for Classification with SVM





    Tutorial for Basic Python Implementation for ANN:

    Walk Through with Python Example in Keras Framework for Classification with Neural Network

    Keras with Tensorflow for ANN
    TensorFlow Example of Classification with Neural Network




    Easy Tutorial on Neural Network Concept
    Easy Tutorial on Multi Layor Neural Network with Softmax Function

    Easy Tutorial on Softmax Function in NN


    Tutorial Sites in R for Classification Steps:

    R Examples:
    Code Examples of Data Cleaning in R:
    Data Analytics (Data Mining) with R
    Multi-Class Classification with Neural Network in R









    Tutorials for Classification with NN:

    scikit-learn: Classification with Neural Network


    Tutorials for Data Preprocessing Methods for NN:

    ML Algorithms -- ANN or SVM Reqire Data Preprocessing:
    Normalization for Any Numerical Attributes
    Binarization (One Hot Encoding) for Any Categorical Attributes

       Tutorial for Data Preprocessing Normalization and One Hot Encoding for ANN or SVM

    Categorical Data transformation Methods with Binarization (One Hot Encoding)

    Data Preprocessing Methods for ANN or SVM


    Tutorial for How to Implement Classification with NN in Python Implementation:

    Walk Through with Basic Example using Panda for Classification with Neural Network **********************

    Easy Tutorial on Neural Network Concept

    Walk Through with Basic Python Example in Keras Framework for Classification with Neural Network

    TensorFlow Example of Classification with Neural Network

    Keras with Tensorflow for ANN





    Easy Tutorial on Multi Layor Neural Network with Softmax Function

    Easy Tutorial on Softmax Function in NN







    How to preprocess Mixed Data Types with Continious and Binary Data Types for SVM
    How to preprocess Mixed Data Types with Continious and Discrete data Types





    Tutorials with Examples:
    Tutorial: Methods to Handle Imbalalnced Class in Python

    Classification Example Codes in Python:
    Examples of Classification: Prediction Model for Hotel Cancellation

    Feature Selection Methods for hotel canllelation classification model in Python




    Examples of Python Implementation for SVM:

    Tutorial: Understanding SVM for Classification

    scikit-learn: Classification with SVM with Code Examples




    Other Good SVM Packages

    Advanced Classification: LIBSVM
    Advanced Classification: SVM Kernel Functions
    Advanced Classification: SVM Kernel Functions


    If you can’t use the scikit-learn package with our own customized weight score, search for the original LibSVM package sites below.
    https://pypi.org/project/libsvm/
    https://github.com/ocampor/libsvm
    https://www.csie.ntu.edu.tw/~cjlin/libsvm/
    https://www.csie.ntu.edu.tw/~cjlin/papers/guide/guide.pdf
    https://stackoverflow.com/questions/17068720/documentation-for-libsvm-in-python


    Paper using SVM:
    Paper Using SVM with Different Kernel Functions for Classification








    Good Data Sets for Classification:

    Health Data Sets for Challenges ****************
    US Government Data Sets ****************





    The Most Important Things to Learn for Classification:

    - Preprocessing methods to transform bigdata to accurate vector form.
    - Understanding the algorithms with the differences in input parameters
    - Understanding the problems, limitations of each ML, and important factors that affect a ML model accuracy and - Solutions to fix the problems and limitations
    For example, Try multiple classifications with SVM and different kernel functions to see which kernel function generates the best model in accuracy.


    Important Problems in Classification

  • Curse of Dimensionality
  • Feature (Selection) Reduction Methods
    - PCA (Principle Compoment Analysis) -- Traditional Statistical Method
    - Correlation Analysis
    - ML Based Recursive Feature Elimination (RFE) -- Advanced ML based Method

    Sklearn: Recursive Feature Elimination (RFE)
    Tutorial: Recursive Feature Elimination (RFE)
    Tutorial: Recursive Feature Elimination (RFECV)



  • Model Overfitting and Underfitting Problem



  • Class Imblance (Minority Class) Problem
  • Handling Imbalanced Data for Classification:
    Learn how to do random sampling with SMOTE for Imbalance Data for Anomaly Detection

    Synthetic Minority Oversampling Technique (SMOTE)

    Imbalanced-Learn Library
    SMOTE for Balancing Data
    SMOTE for Classification
    SMOTE With Selective Synthetic Sample Generation
    Borderline-SMOTE
    Borderline-SMOTE SVM
    Adaptive Synthetic Sampling (ADASYN)



    Tutorials with Examples:

    Handling Imbalanced Data with SMOTE in Python

    Common Sampling Methods to Handle Imbalalnced Class in Python



  • Small Data Set Problems
  • Random Sampling Methods
    (See at the End of the Lecture Notes on Model Evaluation)
    - Bootstrap Sampling (Sampling with Replacement)
    - Stratified Sampling
    - Stratified BootStrap Sampling


    Resampling Methods for Small Sample Problem for Classification:

    - Monte Calro Simulation
    - SVM SMOTE












    10


    Error Estimation and Model Evaluation

    Lecture Notes_14: Lecture Notes (Combined from Both Textbooks) on Error estimation and Validation
    Corrections in Lecture Notes !
    Multi Class Model Evaluation and Comparison: Macro and Micro F1: Precisions/Recall

    Useful Tool for Model Evaluation:

    Model Evaluation and Comparison: Example of ROC with Precision and Recall
    Model Evaluation and Comparison: Macro and Micro F1: Precisions/Recall



    Important Problems to Resolve to Increase Model Accuracy of Classification

  • Curse of Dimensionality

  • Feature (Selection) Reduction Methods
    - PCA (Principle Compoment Analysis) -- Traditional Statistical Method
    - Correlation Analysis
    - ML Model Based Recursive Feature Elimination (RFE)
  • SVM-RFE
  • RF-RFE



  • Model Overfitting and Underfitting Problem



  • Class Imblance (Minority Class) Problem

  • Handling Imbalanced Data for Classification:

    Learn the advanced sampling methods with SMOTE for Imbalance Class Problem for Anomaly Detection

    Synthetic Minority Oversampling Technique (SMOTE)

    Imbalanced Class Learning Library
    SMOTE for Balancing Classes
    SMOTE for Classification
    SMOTE With Selective Synthetic Sample Generation
    Borderline-SMOTE
    Borderline-SMOTE SVM
    Adaptive Synthetic Sampling (ADASYN)


    Handling Imbalanced Data with SMOTE in Python
    Tutorials with Examples:
    Common Sampling Methods to Handle Imbalalnced Class in Python
    Handling Imbalanced Data with SMOTE in Python



  • Small Data Set Problems
  • Random Sampling Methods
    (See at the End of the Lecture Notes on Model Evaluation)
    - Bootstrap Sampling (Sampling with Replacement)
    - Stratified Sampling
    - Stratified BootStrap Sampling


    Resampling Methods for Small Sample Problem for Classification:

    - Monte Calro Simulation
    - SVM SMOTE





    11



    Recommendation System/ Ranking System:

    Lecture Notes on Introduction of Recommendation System
    Lecture Notes on Collaborative Filtering for Recommendation System (Stanford)

    Related Research Papers:

    Papers related to Amazon Collaborative Filtering Algorithm
    Research Paper: Collaborative Filtering for Amazon Recommendation System /a>

    Paper Presentation: Amazon Recommendation System for Collaborative Filtering

    Paper: Recommendation System extended from Collaborative Filtering by Stanford Project Report

    Research Paper: Recommendation System of YouTube with Deep Neural Network


    Recent Research Papers on Recommendation Systems:

    Alicoco large Scale Recommendation System ACM KDD 2021

    Alicoco Large Scale Recommendation System ACM SIGMOD 2020

    Commodity Embedding Alibaba KDD2018

    YouTube Neural Network Recommendations

    Recurrent Neural Network for Netflix Recommendations








    11



    Anomaly Detection

    Lecture Notes_19: Lecture Notes on Data Mining Anomaly Detection (Kumar Textbook)

    Lecture Notes_19_1: Lecture Notes on Data Mining Anomaly Detection

    Lecture Notes_20: Lecture Notes on Data Mining Anomaly Detection by J. Han

    Handling Imbalanced Data for Classification:
    Learn how to do random sampling with SMOTE for Imbalance Data for Anomaly Detection

    Handling Imbalanced Data with SMOTE in Python
    Handling Imbalanced Data
    SMOTE Implementation in R
    SMOTE Implementation in Python
    SMOTE Implementation detail in Python

    Tutorials with Examples:
    Common Sampling Methods to Handle Imbalalnced Class in Python
    Handling Imbalanced Data with SMOTE in Python
    Tutorial: hotel canllelation modelling with knn SMOTE Implementation for Imbalalnced class in Python
    hotel canllelation modelling with SVM for Imbalanced class in Python


    Research Papers on Methods to Handle Imbalanced Data:
    SMOTE: Imbalanced Class Sampling Methods

    Paper: Supervised Anomaly Detection - Classification Methods with Imbalanced data




    ACM KDD Cup Data Sets for Anomaly Detection:

    KDD Cup Data Set for Anomaly Detection

    NSL-KDD Data Sets for Anomaly Detection: Corrected KDD Cup Data Sets

    Kaggle: KDD data set for Anomaly Detection (Use this data for your project on Anomoly Dectection !)
    Data Description for NSL-KDD Data Set for Anomaly Detection
    NSL-KDD Data set Related Research site on AnomalyDetection



    Anomaly Detection Research with AWID Data: **********************

    AWID Data Site: Intrusion Detection in Wireless Network Server Log Data
    Publicaton Site Using AWID Data Set for Intrusion Detection in 802-11 Networks Empirical Evaluation of Threats

    AWID Data Set Available here: AWID: Wireless Network Server Log Data (1 GB zip)
    See More Help on AWID data processing in Lab4 Section !

    Data Collection Paper that describes AWID data set:
    Paper: draft Intrusion Detection in 802-11 Networks Empirical Evaluation of Threats




    Significant Recent Research Papers on Anomaly Detection using AWID Data (For Your Final Project): !!!

    (IEEE Transacton 2018) Deep Abstraction Weighted Feature Extraction and Selection for Dimensionality Reduction

    (2020 ACM Transactions of Intelligent Systems) A Semi-Boosted Nested Model with Sensitivity-based Weighted Binarization for Multi-Domain Network Intrusion Detection

    (KDD 2019) Unsupervised Deep Learning Based AnomalyDetection

    Anomaly Detection in Wireless Network IEEE 2018

    Paper Using SVM with Different Kernel Functions for Anomaly Detection Classification
    Unsupervised Anomaly Detection



    11-12



    UnSupervised Learning ML Algorithms:

    Clustering

    Lectures:

    Lecture Notes_14: Lecture Notes on Clustering Analysis

    Lecture Notes_16: Lecture Notes on Alternative Clustering Analysis (UCLA) Adopted From J. Han

    Lecture Notes_16_1: Lecture Notes on Alternative Clustering Analysis by J. Han

    Lecture Notes_16_2: Lecture Notes on Clustering Clarans by J. Han




    Advanced Clustering Algorithms:
    Leiden Graph Clustering Algorithm


    Leiden Clustering for Community Detection

    Research paper: Spatial Leiden Clustering (BMC 2025)



    scikit-learn Clustering

    Tutorials:

    Lecture Notes_16_2: Tutorial on Clustering Analysis with Mixed Data Types and Visulaization in t-SNE (t-distributed stochastic neighborhood embedding)

    Determining the number of clusters in a data set (From Wiki)

    Tutorial for Clustering with Example Project: Segmentation of Customers by Their Credit Card Usage

    Cluster Visualization:

    Clustering Visulaization: PCA vs t-SNE (t-distributed stochastic neighborhood embedding)






    Hot Spot Analysis for Geospatial Data Analysis

    Lecture:

    Lecture: Hot Spot Analysis in PAI and PEI Metrics


    What is Hotspot Analysis and Z-score and p value based metric


    Wiki on Hot Spot Analysis



    National Institute of Justice (NIJ) Crime HotSpot Analysis Challenging:

    NIJ (National Institute of Justice) Crime Forecasting Challenge Overview



    Research Project Report on PREDICTIVE HOTSPOT MAPPING ANALYSIS (The Paper that Created PAI and PEI Metrics to Evaluate Hot Spot Analysis)

    Documents/Related Materals for NIJ Crime Forecasting Challenges:


    Documentation in Depth for Geospatial Statistics Analysis

    Data Sets:
    NIJ Challenging Data Site: NIJ Crime Location Data Set for Hot Spot Analysis
    NIJ Challenging Data Set: Crime locations in Portland


    For Those Who Want To Apply for NIJ Crime Forecasting Challenges:
    FAQs for NIJ Crime Forecasting Challenges






    GIS Data Visualization API:

  • How to Visualize NIJ GIS Data in QGIS (by Asanka Mananayaka)
  • NIJ GPS data with ArcGIS

  • Useful sites for Geospatial data visualization map using ArcGIS API

    NIJ GPS data with ArcGIS
    Example Project to Guide How to Process Geo Spatial Data for Hot Spot Analysis
    Lab2 GIS Data Processing Example of NIJ Geo Spatial Data Visualization Using ArcGIS Map for Hot Spot Analysis
    Example of NIJ Geo Spatial Data Visualization Using ArcGIS 3D Map for Hot Spot Analysis
    Example of GIS Data Processing in Java Script to Visualize in HTML





    NASA Webserver Log file:
    NASA Webserver Log file Download (NASA HTTP Access Logs) (Scroll Down to the bottom)

    Clustering with NASA HTTP Access Logs See an example project guide




    AWID Site: See Lab4 Section for more Guides
    AWID: Intrusion Detection in Wireless Network Server Log Data
    AWID Data Set Avaliable here:
    AWID: Wireless Network Server Log Data (around 20 GB zip) OR
    AWID: Wireless Network Server Log Data (around 20 GB zip)This site only availabe until Nov 17 !

    Unsupervised Anomaly Detection - Clustering Methods

    Intrusion Detection in 802-11 (AWID) Networks Empirical Evaluation of Threats

    AWID Publication Site: Network Server Log Analysis with Clustering

    NIJ Publication for Hot Spot Analysis


    14



    UnSupervised Learning Algorithms:

    Association Rule Mining

    Lectures:

    Introduction to Association Rule Mining (Kumar Book)

    Association Rule: Apriori Algorithm and Association Rule Mining (J Hans's Book)

    Lecture Notes on Association Rule Generation and Frequent Pattern Tree Algorithm

    Typo Correction on Lecture Notes of Frequent Pattern Tree and Association Rule Generation

    Opimization of Association Rule Mining :Frequent Pattern Tree (Updated in 3rd Eds of J Han)


    Original Paper on Association Rule Mining by IBM in SIGMOD 98


    Text mining using Association Rule Mining:

  • Frequent Pattern Tree in Python
  • Frequent Pattern Tree to Download


  • Association Rule Mining
  • Tutorial for Association Rule mining Using MS Analysis Service
  • Text Mining Examples Using MicroSoft Data Analysis Tool



  • 15



    Deep Learning:


    Good Platforms for Training to Develop Deep Learning Models or Large Language Models: Free GPU Use

  • Google Colab
  • Tutorial: Getting Started with Google Colab
  • Medium by MIT
  • Open AI







  • Basic Review For Multimedia Files:
    Basics on Color Representation: 24 bit RGB Code in 3 Color Channels (One Byte Code (2^8: 0 ~ 255) per Each Channel)
    Basics on Digitizing Images with Resolution
    Basics on Color Image Quantization
    Basics on Digitizing Sound









    Earlier Computer Vision Research

    Face Recognition:
    Open Face Research Site (CMU) on Face Recognition Techniques Using Machine Learning
    Step by step Tutorials on Face Recognition Techniques Using Machine Learning
    Python with Fully Pre-configured VM for Face Recognition Data Analytics
    Simple Tutorial Site for Face Recognition With Deep Learning CNN

    Research Papers on Face Recognition:
    Paper: Deep Face Recognition bt parkhi from KDD 2015
    Paper: Google Embeddings128 Measures from KDD 2015










    Basics Methods to Learn Image BEFORE Deep Learning with CNN:

    For Comparison with CNN:

    Lecture on Basic Notation, Image File as Input with Logistic Regression and intro to Gradient Descent and Neural Network (From DeepLearing.AI by Stanford Coursera)



    Tutorial for Multi-Class Classification with DNN, SoftMax Function and Multinomial Logistic Function for Object Identification

    Before Deep Learning: Simple Example of Object Classification with Neural Network, SoftMax Function and Multinomial Logistic Functions

    Softmax Function

    Softmax Function on Wiki






    Deep Learning for Image Recognition **********************************************

    Lecture Notes:

    Lecture Note: What is a Convolution Function?


    Lecture Notes (with More Descriptions) to Understand Convolution Neural Networks and Deep Learning (From Lecture Note from Connell University)*********************************

    Lecture Notes on Overview of Convolution Neural Networks and Deep Learning with Well Known Deep Learning Architectures (Stanford CS231n) *****************************

    Lecture Notes (with More Descriptions on Optimization Techniques, Transfer Learning, Hyperparameter Setting) on Convolution Neural Networks and Deep Learning with Practical Issues in Learning (From MIT Lecture Note)


    Wiki site Summary on Convolutional Function



    Review Lectures on Deep Learning Building Blocks:

    Lecture Notes_21: Lecture Note on Neural Networks: feedforward

    Advanced Classification: Convolution Neural Networks: Feature Extraction

    Advanced Classification: Convolution Neural Networks: Pooling

    UnSupervised Learning: Convolution Neural Networks: PCA Whitening

    Lecture Notes_23: Lecture Notes on SVM



    Deep Learning Optimization with ResNets (Residual Learning)

    Paper: Deep Residual Learning for Image Recognition

    Wiki on Deep Residual Learning
    Tutorial: Understanding and Visualizing Resnets

    Tutorial: Understanding with Codes for ResNets



    ImageNet Large Scale Visual Recognition Challenge (ILSVRC)







    Tutorial: What is encoder - decoder?








    UNET:

    Tutorial:SUMMARY of UNET

    Tutorial: Understanding UNET

    Tutorial: Understanding UNET







    Current State of the Art Deep Learning for Object Detection

    YOLO Series:

    Object Detection State of Art: YOLOv8
    Object Detection State of Art: YOLOv7 Paper
    Github: Object Detection State of Art: YOLOv7(You need to get an access to this site)

    YOLOv7 Guide

    Object Detection State of Art 2022







    GAN (Generative Adversarial Networks):

    Lecture:
    Lecture on Intro to GAN (Generative Adversarial Networks)

    Tutorial on How to train GAN






    Research Papers: (More Recent Ones to Come here...)

    Paper: ImageNet Classification with Deep Convolutional Neural Networks from KDD 2016
    Paper: Going Deeper with Convolutions from KDD 2015
    Paper: Attention Based with CNN for Question Answering System 2015
    Paper: CoAttention Based with CNN for Question Answering System 2017
    Paper: Deep Residual Learning (2016 Microsoft)




    Image Data Sets for Object Identification with Deep Learning:

    Well Known Deep Learning Image Data Sets

    Image data sets

    CIFAR-100 and CIFAR-10 32 x 32 Image data sets





    Basic Tutorials with Examples of Image Classification to Start With:

    Tutorials for Google Tensorflow for 10 image classification

    Example of a Filter for CNN






    Good Image Classification Tutorial sites:

    Tutorial: getting Started with Google Colab (Python)

    OpenCV Tutorial: Open Source Computer Vision in C (Python or JavaScript Wrapper Available)

    Tutorial Site for Google Tensorflow on Deep Learning with Convolutional Neural Networks

    Python Deeep Learning API: Keras Tutorial

    Tensorflow with Keras Tutorial

    Recurrent Neural Network(RNN) in Tensorflow with Keras Tutorial


    Tutorial on Basic Image Recognition with Classifier

    Tutorials for Google Tensorflow Image Recognition with CNN



    Tutorials for Google Tensorflow with CNN

    Google Analytics and AI

    Google AI Tools

    Google Xception for Image Recognition

    Tutorials for Classes of Linear Algebra in Google Tensorflow

    Tutorials for Image Encoding Decoding in Google Tensorflow








    Summary of Lectures on Deep Learning (From Stanford Class CS229 cheatsheet):

    Deep Learning
    Supervised Learning
    Unsupervised Learning
    Tips and Tricks






    13




    Large Language Models (LLMs)


    Review:

    Overview of Question Answering (QA) System -- Building Artificial Intelligence (Amazon Alexa, Apple Siri, IBM Watson) :

    Question Answering System and Query Analyzer

    Building IBM Watson IBM Watson System (AI):





    Learning Theory of LLMs: Learning a Sequence of a Sentence (Phrase)

    Natural Language Processing (NLP) with Machine Learning




    Problems (Limitations) with Lexicon Based TF-IDF for Text Analysis:

    Phrase (N-Gram Word) Identification
    Negation Handling
    Can NOT Identify Relationships among Terms - Synonyms (Similar Meaning) or Antonyms (Opposite Meaning) of Terms
    Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence - Polysemy Problem !
    Can NOT Identify New Terms, Common Slangs or Changining Relationship Between New Terms (ex: Data Analytics) and Old Terms (ex:Data Mining)
    Can NOT Identify Sarcastic Contexts
    Can NOT Identufy the Order of Words in a Sentence !!!!


    Problems of Learning a Sequence of Tokens in a Sentence/Phrase

    Old Method to Learn Word Sequence: Positioning Index for Phrase Query (Stanford)












    How to Learn from a Large Collection of Document Corpus of Texts/Webpages/Papers/Documents


    Advanced Text Learning Methods with Natural Language Processing and Deep Learning for AI

    Semi-Supervised Learning Methods for Text Analysis:

    Word2Vec (Google):

    Context Aware NLP Methods Using Machine Learning with Semi-Supervised Learning

    Problems of the Lexicon Based Approach for Learning a Sequence of Tokens in a Sentence/Phrase

    Lecture on Earlier Approach: Positioning Index for Phrase Query (Stanford)


  • How to Learn Neighbor Words of Terms(Words) in Sequence in a Phrase or a Sentence
  • Natural Language Processing (NLP) with Word2Vec for Word Embeddings:

    Lectures for Word2Vec:

    Lecture Notes_19_1: Lecture Notes on Natural Language Processing with Deep Learning (Stanford): Word2Vec in Skip Gram Algorithm

    Lecture Notes_19_3: Lecture Notes on NLP (Stanford): More On Word2Vec with Objective Functions

    Stanford Lecture on Natural Language Processing (NLP)



    Lectures on Background of Multinomial Classification

    Softmax Function Definition for Multinomial Classification



    Word2Vec Tutorial With Simple Examples:

    Easy Tutorial for Word2Vec with CBOW and Skipgram Training Model

    Tutorial on Word to Vector (Word2Vec): Skip Gram Model

    Tutorial on Optimization of Word to Vector (Word2Vec) Training








    Original Word2Vec Implementation by Google AI

    Word2Vec Papers:
    Original Word2Vec Papers by Google AI

    Negative Sampling: Optimization Paper for Word2Vec by Google AI





    Extended Word2Vec: Glovec
    Sites for Glove Word Vector by Stanford NLP Team




    Important Word2Vec Papers from Google AI

    Distributed Representations of Word in Vector Space - SkipGram by Google (ACM 2013)
    Distributed Representations of Word and Phrases (Negative Sampling) by Google (ACM 2013)
    Word2Vec: Distributed Representations of Sentences and Documents by Google (ACM 2013)
    ParagraphVec:Distributed Representations of Sentences and Documents by Google (ACM 2014)
    DocumentEmbeddingParaVector by Google (ACM 2015)



    Word2Vec Implementation Sites:

    To get an Executable Binary of word2Vec Model Implementation and Training Data sets:

    Sites for Word To Vector by Google

    Sites for Word To Vector for npmjs
    Sites for Glove Word Vector by Stanford NLP Team
    CNTK: Stanford NLP Sentiment Analysis



    Tutorial Sites for Word2Vec Implementation:

    Tutorials for Word To Vector: Skipgram Model
    Tutorials for Word To Vector with Google Tensorflow
    Tutorials for Word Embeddingd with Google Tensorflow

    Tutorials for Word To Vector and related APIs, Gensim (LDA)
    Tutorials for Bag of Words, Word2Vec and related APIs

    Tensorflow:
    Tensorflow Tutorials
    Tutorials for Word Embeddingd with Google Tensorflow
    Tutorials for Tensorflow with Keras






    Summary of Text Similarity Measuring Methods:

    Text Similarity Measuring Methods





    Data Sets for Document Clustering, Phrase Search or Sentiment Analysis:
    Good NLP Data Sets

    Automatic Frequent Feature(Topic) Discovery in the Hotel Review Data set for Sentiment Analysis

    Hotel Review Data Set (~1.3GB) from Trip Advisor

    You can download preprocessed Wikipedia texts (in XML) here:
    Wikipedia Texts

    Arxiv paper repository to Download
    IMDB Movie Review Repository to Download


    a Good Research Paper on Training Medical Documents with Knowledge Graph (KDD 2019)














    Time Sequence Learning:

    Lectures on RNN/LSTM (Long Short Term Memory)

    Lecture Note: LSTM (Long Short Term Memory)


    Simple Tutorial on Understanding LSTM
    LSTM in wiki

    Lecture Note: More on RNN (Recurrent Neural Network) and LSTM (Long Short Term Memory)

    Lecture Note: GRU (Gated Recurrent Unit)

    Backpropagation Through Time (BPTT) in RNN/LSTM















    BERT: Bidirectional Encoder Representations from Transformers (Google AI)

    NLP Large Deep Learning Model (DLM) for AI: Google BERT Transformer

    Google BERT:

    Leture Note on BERT

    For Advanced: Leture Note on LSTM and Foundation of Large Language Models for Contextualized, Token Based Representation (Harvard Lecture Note)


    Well-known Tutorials on BERT:

    Tutorial: What is Google AI BERT Transformer?

    Introduction to BERT by Mccormickml

    BERT from Original Google AI

    BERT Repo Site to Download


    BERT Papers:
    Original BERT Paper: Attention Is All You Need (Google)

    BERT: Pre-training of Deep Bidirectional Transformers (Google)

    Passage Reranking Using BERT (Google)





    RNN vs LSTM vs GRU vs Transformers



    BERT vs GPT3





    GPTs and ChatGPT from OpenAI:

    GPT3 Paper: LLM are Few Shot Learners OpenAI 2020 Advances in Neural Information Processing Systems ( NeurIPS 2020)

    Tutorial on GPT3 and Training: Few Shot Learners Training with Context

    Tutorial on GPT3 and Applications

    Tutorial on GPT2


    Research Papers at OpenAI

    GPT3 OpenAI

    ChatGPT at OpenAI

    GPT4 at OpenAI




    Stanford Seminar Lecture Series on Transformers











    Overview of Question Answering (QA) System -- Building Artificial Intelligence (Amazon Alexa, Apple Siri, IBM Watson) :

    Question Answering System and Query Analyzer

    Building IBM Watson IBM Watson System (AI):



  • RAG (Retrival Augumented Generation): Document Vector Databases (by Facebook MetaAI)

  • How to Represent Knowledge in Text Documents for ML/AI Algorithm to Learn to Answer?

    Retrieval-Augmented Generation (RAG):

    Lectures:

    Lecture Notes on FAISS and DPR: Intro to RAG for Building a Document Vector Knowledge Database for Retrieval


    Overview of Current RAG Based Approaches and Research Trends with KG Methods for QA Systems









    Recent Research Papers on Large Language Model Training Methods with Context in Prompting :

    Prompting Based Training:
    Research on Prompting Based Training

  • Final Project Guides with Resources and Reserach Paper List on GPT Training Methods with Context Information for Medical Question Answering


    Recent Research Papers on Training Large Language Models (LLM) on Medical Domain:

    BIOGPT
    ChatGPT 3.5 Natively Performs Chain of Thoughts 2023
    Least to Most Prompting 2023
    RankVicuna: Reranking GPT Prompt Model 2023

    Paper Topic Summarization Weak Supervision CMU 2020


    Data Set for Question Answering Training

    32 Million Pubmed Paper Repository Data Sets
    Pubmed Paper Repository Download Site
    Pubmed QA Data Set at HuggingFace
    Pubmed QA

    NLP Data Sets






    Tutorails for BERT :

    Tutorials on BERT Embeddings

    Good Tutorial Site How to Use Pretrained BERT Transformer to get word embeddings

    Install the pytorch interface for BERT by Hugging Face. This library contains interfaces for other pretrained language models like OpenAI’s GPT and GPT-2.
    Huggingface Models, Libraries, Data sets

    Simple Tutorial for Sentiment Analysis with BERT


    Code Tutorial Site to Use Pretrained BERT Transformer:

    Good Tutorial Site to Use Pretrained BERT Transformer

    Good Tutorial to Use Pretrained BERT Transformer for Question Answering System

    Good Tutorial to Use Pretrained BERT Transformer for Domain Specific System

    Good Tutorial Site to Use BERT Transformer for Sentence Classification

    SpaCy Embeddings for BERT Transformer















    Graph Neural Network (GNN)

    Lectures on GNN

    GNN Lecture Notes (Stanford)

    GNN Lectures (Stanford)

    How to Train Knowledge Graph in GNN with Large Language Models

    GNN How to Implement GNN


    Research Paper:

    Graph-Based_Semi-Supervised_Learning_A_Comprehensive_Review


    Graph-based Biomedical Literature Anwering with Knowledge Graph

    Knowledge Graph Embedding Based Question Answering

    BioKG: A Knowledge Graph for Relational Learning On Biological Data












    Reinforcement Learning:

    Lectures

    Lectures on intro to Reinforcement Learning (Stanford)

    Lectures on Reinforcement Learning (Stanford)

    Tutorial on Reinforcement Learning

    Github: Reinforcement Learning




    Applications of Reinforcement Learning







    Stanford Lecture Sites on Related Subjects of AI










  • 16



    Presentation of Projects and Research Papers


    ==> Completion of Homeworks/Labs is required for obtaining a passing grade.

  • This is a tentative scale and
    it could be changed

    Letter
    Grade

    Quality Points

     


      A

    > 93%  

    A: Outstanding (student's performance is genuinely excellent)

      A-

    90% - 93%

     

      B+

    87% - 90%   

     

      B

    82% - 87%

    B: Very Good (student's performance is clearly commendable but not necessarily outstanding)

     

      B-

    80% - 82%

     

     

      C

    75% - 80%

    C: Good (student's performance meets every course requirement and is acceptable; not distinguished)
        D 65%-75% D: Below Average (student's performance fails to meet course objectives and standards)

     

      F

    <65%

    F: Failure (student's performance is unacceptable)

    ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106.

     


    Programming standards

    • Every program must include your name, CSU ID number, Class, Section Number, Hours, the words 'Homework # ...', and a short description of the assignment. For example:
       ' Name: Mark Zuckerberg  
       ' ID: 1234567            
       ' Homework #1            
       ' Description: Computing the average life of a light bulb
    • Every variable should have a meaningful name (this includes function/procedure/subprogram names).
    • Every portion of the program should be as cohesive (single purposed) as possible. This leads to a large number of small functions.
    • Every function (including the main function) should be preceded by a comment indicating its arguments and a description of the transformation it performs.
    • Non-obvious code within a function should be explained.
    • Code should not be over commented.