DSA460/CIS492/593

Data Mining and Machine Learning (4-0-4)

Course Content

  • Class Announcement and Post
  • Class Syllabus
  • Lab Assignments
  • Projects
  • Class Lecture Notes


  • Class Announcement and POST





    You can reach the class webpage from the Big Data Research Lab website as well.
    Big Data Research Lab




    The Final Exam Info:

    For the Active Lecture Notes Links Only for the Final Exam, Access the following site: The rest of links are all disabled.

    the Active Lecture Notes Links Only for the Final Exam

    Hand Written One Page Note (Front and Back) is Allowed to the Final.
    Copy and Pasted Printout of Lecture Notes Are NOT Allowed

    The Exam Format and Rules will be the same as the Midterm.

    Note that You are responsible ONLY for the Lecture Notes and the Subjects Covered in Class

    Question Types:
    6-7 Short Answer Questions on Problem Solving with Given Small Data Sets
    No T/F Questions, No Multiple Choice


    Hand Written One Page Note (Front and Back) is Allowed to the Final.
    Copy and Pasted Printout of Lecture Notes Are NOT Allowed

    The Exam Format and Rules will be the same as the Midterm.

    Subjects to Focus On for the Final

    One question From the Midterm Questions on Classification Algorithms - KNN
    Advanced Classification: NN, Feedforward and Backpropagation Algorithm Calculation
    Clustering Algorithms - K-Mean (Tentative)
    Accuracy Estimation and Validation, Model Evaluation Methods
    Basic Deep Learning Architecture with CNN






    Midterm Info:

    Basically whichever Subjects on the Lecture Notes Covered in Class!

    IMPORTANT WARNING !!
    If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterm and the Final !!

    Midterm Info and The Lecture Note Links Only for Midterm



    Exam 1 will be Tentatively on the Last Week of Feb !

    Subjects to Focus on: 

    Chapter 2, 3 on :
    Basic Stats of Data, Data Preprocessing Methods, Data Transformation Methods, All the Data Proximity Measures, Feature Selection Methods,
    Feature Correlation Measures -- Kai Square, Correlation, Covariance
    Basic Text Analysis Algorithms and Methods in Information Retrieval: TF-IDF Measure, Document as Term Vector  

    See the Topics to Focus on for the Midterm (Exam 1 and Exam 2) here:
    Midterm Links Only (All the links to the rest of the non-related lecture notes are disabled)



    Exam 2 will be Tentatively on the Last Week of March !

    Subjects to Focus on: 

    Classification, Machine Learning Algorithms:
    The Subjects Covered on Machine Learning Algorithms and Their Objective Functions Covered in Class
    Decision Tree, K-NN, Naive Bayes
    Linear Regression, Logistic Regression, ANN Feed Forward, Backpropagation Algorithms, SVM
    Model Evaluation Methods and Common Problems in ML Models for Classification, Data Preprocessing Methods of Each Machine Learning Algorithm for Classification

    Note that You are responsible ONLY for the Lecture Notes and the Subjects that Were Covered in Class



    Question Types:
    6-7 Short Answer Questions on Problem Solving with Given Small Data Sets
    No T/F Questions, No Multiple Choice




    Hand Written One Page Note in Each Side (Only for Equations and Formulas) Is Allowed for the Midterm.







    7. Jan 12, 2026:

    IMPORTANT NOTE !!!

    If You Have a 404 Error in any URL in the Class Lecture Notes under http://eecs.csuohio.edu/~sschung/, You have to replace "http" with "https" to be able to access !




    6. Jan 12, 2026:

    Only the registered students can access the course blackboard.
    If you have a problem with your blackboard access, please contact the register and CSU e-learning Tech Support to resolve the issue !

    Faculty do not control your registration in the CSU Campus Systems and the course blackboard access.

    e-learning CSU Tech Support


    5. Jan 12, 2026:

    It is Required to Attend Every Class !
    Your Class Attendance Will be Checked in Each Class.
    There Will be Random Quizzes to Check the Attendance !
    Only 5 - 10 Mins Will Be Allowed for Each Quiz. Those Who Come in Class Late After 5 Mins, They Will NOT Be Given a Quiz !  


    4. Jan 12, 2026:

    The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class Blackboard !
    Wait For the Lab Submission Links To Be Created on Blackboard with the Deadline to Submit Each Lab




    3. Jan 12, 2026:

    TA Information:

    TA: Abay Suleimenov

    Email: a.suleimenov@vikes.csuohio.edu

    Office Hours: Tues, Thur 10:00AM - 12:00PM or by Appointment for a Zoom meeting
    Make sure to Send him email ahead to let him know you are coming.
    Location: Big Data Lab FH305 or ZOOM Meeting
    ZOOM Meeting ID:
    Passcode:



    If you have questions in Labs or grading your Labs, Send an email to TA to Talk During his TA Office Hours or Schedule a Zoom meeting








    Dr. Chung's Office Hours:

    Tues and Thursday 1:30PM - 3:30PM
    Office: FH 222 or Zoom Meeting
    Email: s.chung@csuohio.edu
    Send Me Email to Set Up a Meeting or a Zoom Meeting.

    Meeting ID: 859 3867 6332
    Zoom Meeting Link




    2. Jan 12, 2026:

    The Output of each lab is your report in Doc file that shows your screen captures of your data processing steps for analytics and the results
    Each of your screen capture must show each step and the result returned by the server in the SAME window in your System to prove that your lab is done correctly in YOUR SYSTEM !!

    Lab Submission:
    1) Submit your Lab in zip file including 1) your lab report in .doc and 2) all the source files, preprocessed input files, outputs on Blackboard for a timestamp as a proof.

    2) If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Heading of the Front Page of Your Report in Bigger and Bold Font !


    1. Jan 12, 2026:

    br> Semester Schedule:
    See University's Official Academic Calendar for the Semester and the Final Exam Schedules

    Final Exam Schedule:
    Tue May 5 4:00p-6:00p





    Exam Info:


    Basically whichever Subjects on the Lecture Notes Covered in Class!

    IMPORTANT WARNING !!
    If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterms and the Final !!

    Midterm Info and The Lecture Note Links Only for the Midterms


    There Will Be 3 Exams (15% Each)

  • Exam 1
  • Exam 1 will be Tentatively on the First Week of March !

    Subjects to Focus on: 

    Chapter 2, 3 on :
    Basic Stats of Data, Data Preprocessing Methods, Data Transformation Methods, All the Data Proximity Measures, Feature Selection Methods,
    Feature Correlation Measures -- Kai Square, Correlation, Covariance
    Basic Text Analysis Algorithms and Methods in Information Retrieval: TF-IDF Measure, Document as Term Vector  

    See the Topics to Focus on for the Midterm (Exam 1 and Exam 2) here:
    Midterm Links Only (All the links to the rest of the non-related lecture notes are disabled)





  • Exam 2
  • Exam 2 will be Tentatively on the First Week of April !

    Subjects to Focus on: 

    Classification, Machine Learning Algorithms:
    The Subjects Covered on Machine Learning Algorithms and Their Objective Functions Covered in Class
    Decision Tree, K-NN, Naive Bayes
    Linear Regression, Logistic Regression, ANN Feed Forward, Backpropagation Algorithms, SVM
    Model Evaluation Methods and Common Problems in ML Models for Classification, Data Preprocessing Methods of Each Machine Learning Algorithm for Classification

    Note that You are responsible ONLY for the Lecture Notes and the Subjects that Were Covered in Class



    Question Types:
    6-7 Short Answer Questions on Problem Solving with Given Small Data Sets
    No T/F Questions, No Multiple Choice




    Hand Written One Page Note in Each Side (Only for Equations and Formula) Is Allowed for the Midterm.







  • Final Exam
  • The Final Exam Info of DSA460/CIS492/CIS593 Data Mining and Machine Learning:

    For the Active Lecture Notes Links Only for the Final Exam, Access the following site: The rest of links are all disabled.

    the Active Lecture Notes Links Only for the Final Exam

    Hand Written One Page Note (Front and Back) is Allowed to the Final.
    Copy and Pasted Printout of Lecture Notes Are NOT Allowed

    The Exam Format and Rules will be the same as the Midterm.

    Subjects to Focus On for the Final

    One question From the Midterm Questions on Classification Algorithms - KNN
    Advanced Classification: NN, Feedforward and Backpropagation Algorithm Calculation, Ensemble Methods
    Clustering Algorithms - K-Mean (Tentative)
    Accuracy Estimation and Validation, Model Evaluation Methods
    Basic Deep Learning Architecture with CNN








    Important Note for the Schedule of the Midterm Exams and Final Exam
    There is always possibility that 2-3 course exams would be clustered in the same day during the Midterm Weeks or Final Exam week. 

    The midterm date is usually decided considering many factors such as the progress of class subjects and the progress of the majority of the students.
    It was announced at least 2 weeks ahead of time.

    I am afraid that it is not possible to change the schedule of the midterms at one person's favor or make anyone take the midterm individually after all the rest of the students in the class had taken the midterm.
    There are many reasons why I can NOT allow it because of serious security bleaches and cheating related concerns.






    DSA 460 Data Mining and Machine Learning is a Core Required Course of the Data Science Degree (BSDS)
    DSA 460 Is Also Offered as CIS492/593 Data Mining and Machine Learning for the CS Majored Students as an Advanced Elective Course

  • Syllabus of DSA 460 Data Mining and Machine Learning (a Core Required Course of the Data Science Degree)

  • Syllabus of CIS492/593/EEC525 Data Mining and Machine Learning (CIS492/593 for the CS Majored Students as an Elective)



  • Lab Assignments



  • Labs:

    First Week Lab0: Choose your System/Tool/Platform to Set Up and Get Used to:

    See the Lab Assignment 0 Section Below to See the Set Up Guide for Python Data Science Platform in Step by Step


    Basic Python Tutorials :

  • Python Tutorial
  • Python Codecademy Tutorial Site


  • Python IDE Deduggers:

    Basic Guide for Python Debugger Pycharm
    Python IDE Spyder
    Python Debugger Spyder



    Python Data Science Platforms:

    There Are Mainly Three Ways to Set up All the Neccessary Data Science Software/Library !

    See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !!
  • Set Up Guide for Python Data Science Platform in Step by Step (by TA Mounika Gandla) ************************************



  • Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system)

    Basic Guide for Python tools for Data Science: jupyter-notebook
    More Basic Guide for Python tools for Data Science: jupyter-notebook


    • Anaconda
    Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms
    Anaconda Tutorials

    • Python Anaconda Tutorial Sites
    Anaconda Tutorial Site
    Anaconda Tutorial Site


    • PyTorch
    PyTorch Site (It can be integrated from Anaconda as well)

    Google Colab:
  • Google Colab

  • • Python Scikit Learn for Common Data Science Tasks
  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing
  • Python Text Processing Libs for Text Analysis
  • Python Numpy Tutorial


  • Text Preprocessing (Natural Language Processing) Library in Python SpaCy:

    Liquistic Modules in Python SpaCy
    Lemmatizer in Python SpaCy
    Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy
    How to Code Liquistic Modules like Lemmatizer in Python SpaCy
    Python Example for Basic Text Processing

    • Python sklearn.cluster Python Sklearn Clustering






    There is Another Data Science Platform in R if You Choose to Learn (SAS Data Miner or MS Data Tool Have an Integrated R Platform)

    Basic R Tutorials :

  • R Studio basic Tutorial
  • R Basic Online Lecture

  • R Manuals

  • R Tutorial with Examples with R Stat Tool
  • Note that Examples in this tutorials may not the final correct output for Lab1 !
  • Examples of Basic R Stat Tool with Helpful References

  • Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study by Nick White (Now in FaceBook)






    Useful Machine Learning Tutorial Sites:

    Keras for Image Processing/Text Processing with Deep Learning

    Google Colab for Fast Machine Learning Execution in GPU


    Good Data Science and AI Platforms:
  • Open AI
  • Google Colab
  • Tutorial: Getting Started with Google Colab
  • Medium by MIT




  • For Your Own Study
  • Coursera Machine Learning by Stanford



  • Lab Submission Instructions:


    The Output of each lab is your report in Doc file that shows your screen captures of your system/tool/platform setting/configurations, your data processing steps for analytics and the results
    Each of your screen capture must show each step and the result returned by the server in the SAME window in your System to prove that your lab is done correctly in YOUR SYSTEM !!

    1. Submit (on Blackboard) your Lab in a zip file including 1) your lab report in .doc and 2) all the source/scipt files, preprocessed input files, outputs on Blackboard for a timestamp as a proof.
    2. If you did an Extra Credit Lab, Make a Note on the COVER of Your Lab Clearly !


    The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard !
    You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission

    If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font !

    Always Follow the Deadline of Each Lab Assigned on the Class Blackboard.
    The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester.

    Please Identify Your Course When You Ask Me in Email !



  • The Output of each lab is your Lab Report in Doc file that shows your screen captures of each of your executions with Your Outputs

  • Your Report in Doc file should include all the platform set up procedures, the execution steps, and copy of each source code files
  • Each of your screen capture must show your results returned by your experiment to prove that you have done the lab correctly !!



  • 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof.

    2. IMPORTANT NOTE:

    Your Screen Captures in Your Lab Report Should Show Your Own System Name to Prove That Your Lab Was Done by YOU on Your Computer.

  • TA Instructions on How to Create Your Lab Report
  • How to Create Table Contents in a word Doc File for Your Lab Report
  • Example of Output of the Execution Steps for Labs

  • 3. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font !

    4. Always Follow the Deadline of Each Lab Assigned on the Class Blackboard !!
    The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester.






    Lab Assignment 0:


    Set Up Your Python Data Science Platform for Labs and Final Projects by the End of the First Weekend !

    There Are Mainly Three Ways to Set up All the Necessary Data Science Software/Library !

    See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !!
  • Set Up Guide for Python Data Science Platform in Step by Step (by TA Mounika Gandla) ************************************









  • (Not for This Semester !)
    PreLab 0 : Preliminary Analysis of the Computer Game Sales Data (Not for this Semester !)


    With the Given Table - Global Video Game Sales data (below) obtained from the Data Warehouse of a Game Retailor Company,
    Analyze the Game Sales Data to Identify Any Trend in the Total Sales of the Game During Last 10 Years. You need to find the following:

    Which Computer Game Genres, Platforms Will Be the Best to Focus to Invest/Develop for Next Year (in this data set next year is 2017) Globally and For Each Region.

    To be able to predict, you need to analyze sales data to provide any evidence of your prediction:

    Hint: 1. See total sales by Genres, Platforms respectively per each year during last 10 years.

    2. Save each analyzed result into a CSV file (or another excel sheet) for further processing or any visualization tool or Excel to Visualize in Graph to Compare total sales of each Genres, Platform during last 10 years to find any useful intelligence to Predict

    3. Report the facts found from your visualization of each analysis result and discuss why you predict that the specific Genres and Plaform are best to invest in for the next year.

    You may use Any Tool of Your Choice for analysis and visulization -- Even Excel should be enough for this task to analyze and visualize the analyzed results.


    Hint: One easy way to create analyzed data is that you can do all the calculation per each year and each genre , (each year and each platform as well).
    To create a graph, you need each aggregated total per each year and each genre, (each year and each platform as well), saved in one file (this can be done by storing the SQL result in a table), then saved the table as an excel or csv file for a graph.

    Analyze and Visualize the Computer Game Sales Data of a Company below to Predict which Genre and Platform of Computer Games Will Be the Best Sellers Next Year.
    The Predicted Genre and Platform Can Be Provided as Decision Supporting Evidences for the Company to Make a Correct Decision to Invest for Next Year (in this data set, next year is 2017) Globally as well as Each Region.

    You may use any tool of your choice -- Even Excel should be enough for this task to analyze and visualize the analyzed results.

  • Global Video Game Sales Data








  • Although the database skill is NOT required for this course, it will be very useful for handling big data sets.
    For Those Who want to Use a Database Server for Initial Data Handling

    Installation Guides for MySql Server:

    MySql Download
    MySql Download and Installation
    MySql Tutorial
    How to Create MySQL Database


    Microsoft SQL Server:
    See the Lab Sections of CIS 430/530 for Microsoft SQL Server Installation Guides



    Useful SQL Review to Get to Know Your Data Features (Columns)

  • Examples of LOJ with Group By and Aggregate Functions
  • Example Outputs of LOJ with Group By and Aggregate Functions









  • REQUIRED LAB ASSIGNMENTS:


    The Submission Link and the Deadline of Each Lab Will Be Posted on the Class Blackboard !


    Lab Assignment 1:


  • Lab Assignment 1 on Data Cleaning, Preprocessing and Transformation, and Similarity Measure, Similarity Matrix

  • Lab1_1 - Part 1 and Part 2 on Preprocessing and Transformation: Due By the End of the Third Week
    Lab1_2 - Part 3 and Part 4 on Similarity Measure and Building Similarity Matrix Due By the End of the Fourth Week


    Input Data File for Lab1:
  • vTargetMailCustomer Added Outliers Income & Age (csv) Customer Profile Data Set
  • Lab1 Small Data Set (Mouse Serum Data) for Null Replacement (csv) (Three Columns Marked for Null Replacement)


  • TA Instruction on How to Generate Lab1 Report


  • For Part 1 and 2:

    For each selected feature, identify ALL the required data preprocessing methods like Normalization, Discretization, Binarization, and more based on the feature data properties. They must be done along with other preprocessing methods.

    Examples of Lab1 Report of Data Preprocessing and Data Similarity Measures

    Example1 of Lab1 Output: Part 1
    Example of Lab1 Output for the Final Transformed Features

    Note that these Examples Here May NOT neccessarily all correct ! They just show How Lab1 can be done as an example. Do NOT blindly follow !

    For example, EnglishEducation should be Transformed as Categorical or Ordinal?
    If it is transformed as Categorical, One-hot-encoding is a correct transformation.
    However, if it is Ordinal, then one hot encoding is not a correct transformation for this column.



    For Part 3:

  • Final Selected Feature Set

  • NOTE !!!
    You Are NOT Supposed to Use any Python Libs to Calculate Similarity Measures in Part 3 of Lab 1.
    You have to write Scripts/Programs to Compute each Similarity Measure.
    If you use the built-in lib to build a Similarity Measures, which is one line of code, you will get 0 for the part.




    Suggested Platforms to Use for Lab1:

  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing

  • Anaconda

    Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms
    Anaconda Tutorials


  • FAQs on Lab 1

  • FAQs for Lab1:

    Q:
    For Part 1 of Lab 1, is it acceptable to complete the data processing using Excel? I wanted to confirm whether submitting the work in Excel is appropriate, or if you would prefer the responses to be documented in a Word file instead.

    A:
    It was instructed to set up to use one of the Data Science platforms introduced in class as Lab0 and create a Lab report in doc file with the screenshots showing the intermediate results.
    Q:
    Additionally, for Lab 2, should we use the dataset generated from our random sampling, or the dataset that was provided in the lab materials as the input?

    A:
    The Lab1 specification clearly mentioned at the beginning that use a full data set for the tasks in Part 2 as below:
     
    Part 2: Data Preprocessing and Transformation Use all the data rows (~= 18000 rows) with the selected features as an input file to apply the tasks below, do not perform each task on the smaller data set that you got from your random sampling.
    

    Q:
    In Part 2 of Lab 1, it says we need to replace the Null values. What are we replacing them with? Also, later in part 2, it says that we need to replace the outliers. I am also confused what we are replacing the outliers with. A:
    The Null value replacement methods to determine the value to replace Nulls were covered in the class lecture notes at:
    Null Replacement Methods
    Depending on the feature column type, for example, Mean, Median, Min, or Weighted Mean would be a representative value to replace Null with.

    For an Outlier replacement, 1.5IQR to calculate the upper extreme and low extreme to identify outliers and
    replace them with the lower extreme, upper extreme in each side as explained at:
    1.5IQR Outlier Replacement Method


  • Tutorial on Data Preprocessing: One Hot Encoder
  • Data Normalization and Binarization Encoding Example for Neural Net classifier






  • Lab Assingment 2:

  • Lab Assignment 2 on Text Analysis with TF-IDF

  • FAQs for Lab Assignment 2

  • Note That Lab2 Requires Stemming/Lemmarizaion and Bi-Gram and Tri-Gram Handling

  • Common NLP Preprocessing Tasks to Be Done for TextAnalysis
  • Input File 10 State Union Address (xlsx file)
  • Input File All State Union Address (CSV file)


  • You can use any avaialble API for Preprocessing

  • Set Up Beautiful Soup for Web Scraping for Text Analysis

  • Some Trouble Shooting Tips:
    For Lab 2, it requires the python library 'clean-text'
    The full command to install this dependency is:
    pip install clean-text
    If you do wish to rerun it and make sure that it works, you would have to uninstall the old one first as they use the same module name.
    Having both installed at the same time will result in Python using the incorrect one.


    NOTE !!!
    You Are Not Supposed to Use any Python Libs to Calculate a Cosine Similarity for Lab 2.
    You have to write a script to compute and build a Similarity Matrix.
    If you use the built-in lib to build a Cosine Similarity, you will get 0 for Labs.








    Lab Assingment 3 on Classifier

  • Lab Assignment 3 on Classification with Machine Learning



  • IMPORTANT NOTES for LAB3:

    1. Use the Best Selected Features Given Below instead of Your Own Selected Feature Set in Lab1 :
  • Target Mail Selected Feature Set

  • IMPORTANT !!: In your Lab report, Make sure to Show the final transformed training set file and all the attribute values for the first two objects
    in your training set that was used for your classifiers.


    2. 4 Classifiers for Lab3 as below:

  • Classifiers from the Probabilistic Approach Based ML Algorithms
  • 1) Decision Tree 2) Bayesian 3) An Ensemble Method: Random Forest 4) K-NN (Distance Based)


    3. Grading is Based on Your Best Accuracy with the best input parameters you identified from Your Experiments


    Tutorial with Codes Examples for Implementation of Classification with ML Algorithms

    Example of Classification with DT

    Example 2 of Classification with DT

    Example of Classification with ANN and SVM








    Lab Assingment 4 on ANN:

    Lab 4_1 on Calculation of ANN Backpropagation:
  • Lab Assignment 4_1 on Calculation of ANN Backpropagation (this is a pencile-and-paper assignment)


  • Lab4_1 - ANN Calculation in Pencil and Paper — Required for All

    Lab4_2 - Classification with ANN - Required for All for All

    Extra Credit Lab4_3 (100 %):
    Lab4_3 - Implementation of ANN Algorithm in Feedforward and Backpropagation in Matrix Operations in Linear Algebra either in Python Numpy or MatLab for Classification




    Lab 4_2 on Classification with ANN and SVM:

  • Lab Assignment 4-2 on Classification with ANN and SVM

  • Use the Same Data Set Given for Lab3 for model comparison

    Requirements for Lab4_2:

    1. Experiment Your Classification with two Numerical Approach Based MLs: ANN and SVM

    2. For Each ML, Find the Best Hyperparameters (input parameters) to Find the Best Performing model
    With ANN, Repeat training to find the optimal hyperparameters (Learning rate, the Numer of Hidden Layer) for the best fit(model) for the given data
    With SVM, Repeat training to find the optimal hyperparameters (Kernel Function) for the best fit(model) for the given data


    Data Set for Lab4_2





    Extra Credit Lab4_3:
  • Lab Assignment 4_3 on Implementation of ANN Backpropagation Algorithm


  • Data Set for Lab4_3:
    Patient Data Set for Breast Cancer Prediction for Binary Classification to Predict whether the Patient has a Breast Cancer or not
  • Health Data Set: Patient Data for Breast Cancer Prediction for Binary Classification

  • See the Lecture Note Section for Implementation of ANN for this

    Implement Feedforward and Back Propagation of NN for Classification

    Step by Step Tutorial of Backpropagation with an example

    Lecture Note on Step by Step Matrix Computation of Backpropagation with an Example ************************

    More Detail on Neural Network with Matrix Implementation of Backpropagation Algorithm (From DeepLearing.AI by Stanford Coursera)



    Matrix Implementaion:
    Standard Notations for Deep Learning

    Sample Codes for Implemenation of Forward and Back Propagation of NN Using Matrix Operations in Python









    Lab Assingment 5 on Clustering:


  • Lab Assignment 5 on Clustering
  • For each Clustering result in your experiment, Apply any method discussed in the Lecture notes (Ward's method, Silhouette score, Elbow method, Entrophy/Purity etc) to Measure the quality of the Clustering result.
  • FAQs for Lab 4 on Clustering
  • Data Preprocessing Issues for Clustering
  • Tutorial on Advanced Clustering Analysis with Mixed Data Types


  • Some Suggested Platforms for Clustering

    scikit-learn Clustering



    Although Clustering in general should work on data as multidimensional vectors,
    For NIJ Data For Lab4, you can work on NIJ Challeneges to identify Hot Spots for Crime Location - GPS data and Crime Category instead of handling data as multidimensional object for Clustering algorithms
    Cluster on the Crime Loactions (X and Y coordinates - transform them to GPS coordinates) with K-Mean and DBCSAN and Cluster the Crime locations for each Crime Category
    but Add PAI or PEI Analysis for Hot Spot Analysis with Changing Parameters

    National Institute of Justice (NIJ) Crime HotSpot Analysis Challenging:

    NIJ (National Institute of Justice) Crime Forecasting Challenge: data description and overview


  • NIJ (National Institute of Justice) Challenge Site

  • NIJ Challenging Data Published in 2017
  • More Recent Data Sets in NIJ Challenging Data Site

  • PAI: How to Measure HotSpot Cluster of NIJ GIS Data in PAI
  • Tutorial: How to Measure HotSpot Clusters of NIJ GIS Data in PAI
  • How to Visulaize NIJ GIS Data in QGIS by Asanka Mananayaka
  • How to Visulaize NIJ GPS data with ArcGIS



  • PROJECT



    Project:



  • Final Project Ideas and FAQs



    You Can Choose Your Own Data Set to Extend Lab3 to Design Your own Classification for Final Group Project !

    Data Set Repository for Classification:

    Kaggle Data Set Repository
    UCI Data Set Repository for Classification
    Health Data Set Repository for Classification
    Socrata Open Data API






    Important Dates (Tentative) for Final Group Project
    (The Exact Deadlines will be Posted on the Blackboard)

    Group Project Proposal Due by April 3rd !
    Group Project Status Report Due by April 17 !
    Group Project Presentation Either on April 28 or 30 !
    Final Group Project Report Due By Friday May 1 !





  • Task 1: Group Project Proposal Due by April 14 !

    Submit Your Group Proposal (in Minimum 3 Pages) on BlackBoard

    Group Proposal Should Include:

    1. Data Description, Data Size, Data Collection Plan (if needed),
    2. Systems/Tools to Use,
    3. Data Preprocessing Methods,
    4. Data Analytic Goal with Plan of Evaluation (Design of Your Experiment) in Detail




  • Task 2: Group Project Status Report ! Due By April 20

    Your Project Status Report Should Show the Following Tasks Done:

    1. Platform Setting/Configuration Procedure if it is new,
    2. Your Data Contents, Selected Feature Description,
    3. Data Preprocessing Steps and the intermediate Outputs




  • Task 3: Project Presentation Starts on Last Week of the Semester !
  • Read the Instructions of Project Presentation Here ! (This Google Project Presentation Scheduler Will Be Sent To Your CSU Email !)

    Your Project Presentation Should Have Contents on:

    1. Data Description, Data Size, Data Collection Method(if needed),
    2. Goal of Your Project for Data Analytics or Data Mining Application (AI)
    2. Systems/Tools to Used,
    3. Feature Selection Method and The Final Features Selected
    4. Data Preprocessing Methods and Intermediate Results, Final Training Set and Test Set Description
    5. Data Analytic Design (of Your Experiment) in Detail
    Show the Source Codes for Each Below:
    5_1 Preprocessing and Transformation Methods Done
    5_2 Size of Training and Test Set
    5_3 How Many Classes in Your Classification and Class Distribution in Size
    5_4 How Training Is Done
    5_5 Model Evaluation Metrics Used
    5_6 Results in Accuracy and Comparisons
    6. Challenges/Problems Encountered - How to Resolve Them
    7. Evaluation Results and Visualization of the Results




    For the Class Size with < 25 Students, 1-2 Person Group Is Allowed
    For the Large Class with > 35 Students, Any Group Size in 1 - 4 Person Group Are Allowed
    Note that Large Groups (3-4 Person Group) Shoud Complete a Bigger Project !


    One Submission Per Group Is Required !
    EACH Memeber Name and ID SHOULD BE LISTED in the COVER !
    List Your First and Last Name Only ! Exactly as Appeared in the CSU CampusNet. DO NOT Use Your Middle Name.

    If We Can't Find You by Your Name appeared on Your Project Report and Presentation Schedule. Your Project will be Considered as Missing with 0.

     



    Task 4:
    Final Group Project Report Submission Instructions (By the End of Friday of Your Presentation Week):

    Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week !

    Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx !

    If your data file is too big to upload, Submit your zip file with your Data file on your ONE drive and Send email to me and TA attached the data file on One Drive.




    Submit a Zip file that includes:

    1. All of your presentation slides (both in .ppt and .pdf) and
    2. Your Group Final Project Report (in doc) that shows with:
    1.Platform/System Set up Procedures/Instructions,
    2. Evaluation Results and Visualization of the Results
    3. Executions Steps, all the source codes/scripts, all the intermediate outputs, and final output files
    4. Include the Problems/Error Encountered and Your Resolutions in Your Report




    One Submission Per Group Required.

    Submit a Zip File that Includes All the required Source Files, Input, Output Files, and Final Report (in Doc file) and Presentation Slides (in pptx).

    Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files.

    The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results.

    If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes from the Web.







    Good Data Sets for Classification:

    Health Data Sets for Challenges ****************
    US Government Data Sets ****************


    More Project List and Data Sets:

    Sentiment Analysis of Online Reviews/Social Media data:
  • Trip Advisor Hotel Review Data Set (from Carnegie Mellon Data set)
  • Amazon Product Review Data Set (from UCSD Research Site) Request data sets on category: Mobil/Cell Phone Reviews Or Electronics
  • Good Conference Site List to Search Research Papers on Review Analysis (From UCSD Research Site)
  • IMDB Moview Review data sets
  • NLP Text Data Set Repository








  • Senior Design Projects on Big Data and AI:

    2022 - 2024:
    Big Data and AI Projects:
  • Intellestate: Intelligent Real Estate Property Recommendation System: The Candidate of Best Senior Project in Engineering College of 2023) (Created From CIS430, CIS492/593 Big Data, and CIS408)
  • Sign to Speech: Sign Language Translating Glove (Created From CIS430 and CIS492/593 Data Mining)
  • Infectuous Disease Predictor (Created From CIS430, CIS492/593 Big Data, and CIS408)
  • Social Media Opinion Analysis for Congress Bills (Created From CIS430, CIS492/593 Big Data, and CIS408)


  • 2020 - 2021:
    Big Data and AI Projects:

  • Social Media Opinion Analysis System for Prediction of 2020 Presidential Election (The Candidate of Best Senior Project in Engineering College of 2021) (Created From CIS430, CIS492/593 Big Data, and CIS408)
  • 2021 Senior Project: Stock Market Analysis Service System (Created From CIS430, CIS492/593 Big Data, and CIS408)
  • Intelligent Infectious Disease Tracking System (Created From CIS430, CIS492/593 Big Data, and CIS408)
  • Product Review Sentiment Analysis System (Created From CIS430 and CIS492/593 Big Data)


  • 2019 - 2017:
    Big Data and Data Science Projects:

  • The Candidate of Best Senior Project in Engineering College of 2019 by Joel Stell et. al. (Created From CIS430, CIS492/593 Big Data, and CIS408)
  • Web Search Engine (Google Like) over Research Paper Repository by Nick McCoy, et al (Created From CIS430, CIS492/593 Big Data, and CIS408)
  • The Best CS Senior Project Winner of 2017 (Created From CIS430, CIS492/593 Big Data, and CIS408) by Mike D'Arcy and Utkarsh Patel












    For the Graduate Courses CIS593 and EEC525 ****************************

    More Advanced Topics for Final Project

    Implement Your Own (Deep) Neural Network Architecture Using Google Tensorflow or MS CNTK or any of your choice for Text Analysis Task or Image Processing: Extra Credit !



    If you want to work on Neural Network, then download source codes of skipgram (the first paper) from one of those sites below to learn how to build Skipgarm with NN then implement paragraph vector as document vector in the second paper.
    Look at the first research project for guide for this project at:
  • Project: Implemnting Paragraph Vector using Word2Vec

    Then, let's read the next two papers to understand the source codes for you to download and do the experiment with real data set.
  • Word2Vec Research Paper from Google
  • Paragraph Vector for Sentences and Document from Google 2014

    Then check these two sites where you can download all the source codes to start an experiment with real data sets.
  • Word2Vec to download (google site)
  • Word2Vec to download with npmjs
    Data set to train to generate word2Vec and Paragraph Vector:
    You can choose any webpage set or papers as data but it should be at least 200,000 documents to train
    There is a preprocessed wiki page data set is available in the Stanford NLP site that can be used as your training data set. You can make this as your project if you want.


    The Most Recent and Most Superior Word Vector: BERT -- See Advanced Text Analysis Section in Class Lecture Notes for more on BERT

    BERT from Google AI

    Tutorials on BERT from Google


    Good Tutorial Sites to Learn How to Use Pretrained BERT:

    SpaCy Embeddings for BERT Transformer

    Good Tutorial Site to Use Pretrained BERT Transformer

    Good Tutorial to Use Pretrained BERT Transformer for Question Answering System

    Good Tutorial to Use Pretrained BERT Transformer for Domain Specific System

    Good Tutorial Site to Use BERT Transformer for Sentence Classification



    Research Paper: Passage Reranking Using BERT from Google


    NSL-KDD data set for AnomalyDetection:

    Kaggle: KDD data set for AnomalyDetection
    Data set Related Research site on AnomalyDetection


    For NIJ Challenging: Hot spot Analysis: See Clustering Section for more detailed information

    See More Info for NIJ Challegeing in the Clustering Section

    Download New data sets and Submission files of Winning Teams below to Learn from the NIJ (National Institute of Justice) Challenege site (Scroll down to see the links)
    Find Out How to Measure Hot Spots in score type PAI and PEI* for Every Category, Crime Type, time Frame.
  • NIJ Challenging with Hot Spot Analysis
  • NIJ New Data Sets
  • NIJ Challenging Data from 2017
  • NIJ Related Articles/Papers
  • Tutorial: How to Measure HotSpot Cluster of NIJ GIS Data in PAI
  • How to Measure Cluster of NIJ GIS Data
  • Documenetation in Depth: How to Approach for Geospatial Data Analysis
  • Examples of Solutions: How to Use Challenges to Find Solutions

  • GIS Data Visualization API:
  • How to Visulaize NIJ GIS Data in QGISby Asanka Mananayaka
  • NIJ GPS data with ArcGIS


  • List of Top Journals and Conferences in Data Mining and Information Processing

  • R Tutorial for Project
  • Project Guide with an Example: Data Mining over LinkedIn Data
  • Project Guide with an Example on Twitter Text Mining
  • Presentation Schedule and List of What to Present in Your Presentation
  • How to read and Present a Research Paper





  • Text Analysis Using Classification
    Text Mining on Twitter Data on ISIS terrorists group and the fan groups of ISIS Santosh Tankala, Vishnu Vishnuteja Thummanapelli, and Akhi Reddy Laxmanagari
    Tutorial: How To get Twitter Data

    Another Projects:

    Sentiment Analysis of Online Reviews:
  • IMDB Moview Review data sets
  • Trip Advisor Hotel Review Data Set (from Carnegie Mellon Data set)
  • Amazon Product Review Data Set (from UCSD Research Site)













  • Textbook Search Site at CSU the Library Site

    Web Accessible Textbook at the CSU Library Site


  • Class Lecture Notes with Tentative Schedule

    Class Chapter / Topic / Specific Objectives / Activities
    1

    Introduction to Big Data Analytics:

    Lecture Notes_1_1: Simple Overview of Big Data Analytics

    Lecture Notes_1_2: Introduction to Big Data, Big Data Processing and Big Data Analytics

    Lecture Notes_1_3: Big Data, Data Scientist and Research on Big Data Analytics at CSU


    1-4


    What is Data Mining?

    Chapter 1: Overview of Data Mining: What is Data Mining and Data Mining Steps


    Basic Data Structures and Data Properties in Data Mining:

    Chapter 2_1: What is Data Structure and Properties of DATA in Data Mining? (Kumar's Chapter 2_1)*****************

    Chapter 2_1: Your DATA in Properties and Data Structures for Data Mining (Part 1) (J Han's Chapter 2_1)



    Example of Data Preprocessing and Transformation Steps For a Data Mining with Classification


    Data Preprocessing Steps: 1. Feature Selection, 2. Preprocessing:Cleaning, 3. Exploring to Get to Know Data, 4. Reduction/Integration, 5.Transformation, 6. Correct Measuring


    1. Feature Selection Methods: (To Be Covered at the end)





    2. Data Cleaning:

    Chapter2_1: Data Quality and Data Cleaning (Kumar's Chapter2_1)

    Chapter3_1: Data Cleaning for Null and Outliers in Data Preprocessing (J Han's Chapter3_1)*****************




    2_1: Getting to Know Your Data with Basic Stats for Cleaning


    Chapter 2_2 : Knowing Your Data for Data Disperse/Distribution (Basic Statistics Review) - Part 1 (J Han's Chapter2_2)***************

    Grouped Data Calculation


    Measure and Visualization of Data Disperse (Distribution) of Each Feature in Data Set:

  • Box Plot
  • Histogram
  • Z-Score





  • 2-2: Data Cleaning:

    - Null Value Replacement Methods:

  • Replace with Mean/Median/Min/Weighted Mean

  • Repalce with the Most Probable Vaue Inferenced from Probability Distribustion of each Value in each Feature


    - Outlier Removal/Replacement Methods:

  • 1.5IQR Method
  • Z-Score Method


  • Tutorials with Examples:
    Simple Examples of 1.5IQR for Outlier Removal Method with Box Plot per each Column

    Simple Examples of Z-Score

    Common Outlier Removal Methods for Each Column: 1.5IQR, Z-score, Percentile

    Common Outlier Removal Methods for Each Column: 1.5 IQR, Histogram, Z-Score, Inference with Regression

    1.5 IQR for Outlier Removal with Box Plot per each Column





    3. More Getting to Know Your Data

    Chapter 3: Overview of Data Exploration (Kumar's Chapter 3)








    4. Data Integraton and Reduction :

    Chapter 2_3: Intro to Data Integration, Reduction Methods (Object Reduction) - Aggregation, Discretization, and Sampling Methods (Kumar's Chapter 2_3)

    Data Reduction (To Reduce the Number of Distinct Vaues) with Discretization Methods: Binning, Histogram (From the slides of J Han's Chapter 3_3)




    Sampling Methods for Data Reduction:

    Bootstrap Sampling with Replacement
    Stratified Sampling with Weighted Mean


    Oversampling (Resampling) Methods for a Small Data Set:

    Monte Carlo Sampling for Each Feature
    Stratified Bootstrap Sampling
















    5. Data Transformation: Normalization, Discretization, Binarization

    Chapter 3_4: Data Transformation - Normalization, Discretization (J Han's Chapter 3_4) *************************

    Chapter 2_2: Data Transformation Basics (Kumar Chap2_2)



    Common Data Transformation Methods for Data Measures:

  • Binarization (Hot Encoding) Method for Categorical (Norminal) Data Encoding
  • Discretization for Data Reduction for Numeric Data with Too Many Unique values
  • Normalization/Standardization for Numeric Data


  • Correct Transformation Methods for ML Algorithms:

    Required Data Processing for ML Algorithms: Decision Tree or Random Forest (Probablistic Based ML Algorithms):

  • Discretization for Continuous or Integer Attributes with too many distinct values

  • Example of Equal Frequency (Depth) Discretization For Classification




    Reqired Data Preprocessing for ML Algorithms ANN or SVM (Numercal Approach):
  • Normalization for Any Numerical Attributes

  • Binarization (One Hot Encoding) for Any Categorical Attributes



  • Binarization (One Hot Encoding) for Categorical Features :

    Examples of Binarization (One Hot Encoding) for Categorical (Norminal) Data Transformation

    Examples of Data Transformation Methods: Normalization for Numeric Data and Binarization (One Hot Encoding) for Categorical Data for Neural Network




      


    How to Transform TimeStamp:

    Examples of Preprocessing DateTime as Features for Classification: Appointment Cancellation Prediction

    Transformation Methods for Date Time Data Type

    Transformation Methods for Delta Time from Date Time Data Type


     




     

    6. Data Measures:

    Chapter 2_4: Data Measures (Kumar's Chapter 2_2)********************

    Chapter_2_5: Knowing Your Data (Part2): Measuring Data Proximity Per Data Properties (J Han's Chapter2_3)*********************




    Correlation Measure:

    Chapter 3_1: Correlation Analysis for Feature Reduction and Feature Correlation (J Han's Chapter3_3)************************

    Expected Value as Weighted Mean of Random Variable


    Lecture on Covariance Matrix for PCA


    Scores and Ranking









    Chapter 3_4: Correlation Analysis (J Han's Chapter 3_2)



    Lectures on Data Reduction and Dimentionality Reduction Methods:

    Dimensionality Reduction Methods for Feature Selection: (Will Be Covered More Later)

    Lecture Notes with Examples of Dimensionality Reduction

    Chapter 2_4: Data Reduction Methods and Feature Selection Methods (Kumar's Chap2_4)









    Tutorial for Feature Selection Methods:



    scikit-learn Feature Selection Methods




    4-5



    Unstructured Text Processing for Text Analytics

    Lexicon Based Approach with Document(Text) as Bag of Words Model:

    Lecture Notes_18: Lecture Notes on IR TF-IDF for Text/Web Mining ******************

    Lecture Notes_18: Lecture Notes on Inverted Index


    Example of Building Inverted Index on State Union Addresses for Text Analysis



  • Example Project on Twitter Text Mining and Sentiment Analysis

  • Sample Projects on Sentiment Analysis with Lexicon Based Bag of Words Model

    Basic Another Sentiment Analysis with Yelp Review Data


    Common Natural Language Processing (NLP) Techniques for Text Preprocessing Tasks:

  • Common NLP Preprocessing Tasks to Be Done for Bag of Words Model with TF-IDF Scorng Algorithm for Text Analysis

  • When There is No Inverted Index Built in the Collection of the Documents for TF-IDF Based Text Anaysis,
    Removing Stop Words from Stop Word List Could Be a Partial Alternative (Not the Same) for the Probelm Resolved by Inverted Document Frequency

    However, Note that You Have to Decide Which Text Cleaning Preprocessing Should Be Applied or Should Not Be Applied in Your Text Processing Data Pipeline Depending on Your Goal of Text Analysis

    For TF-IDF Based/Lexicon Based Text Vectorization, You Can Apply All the Text Cleaning Preprocessing listed above.
    On the other hand, for the Rest of the NLP Parsers(API) such as POS Tagger, NER Tagger, You Should Not Remove Each Sentence Strcuture, Should Not Apply Some of the Common NLP Cleaning Tasks Such as Stop Word Removal Which Will Destroy a Sentence.

    Note that Removing Stop Words is Only for Lexicon (TF-IDF) Based Text Analysis. Removing Stop Words Could Have a Similar Effect (but Not Equivalent to) as Inverted Document Frequency(IDF) when You Cannot Measure IDF.
    However, It May NOT Be Good For Context Aware (Sentence Level) Information Extraction Methods like POS Tagger and NER Parser, or Word Embedding based Word2Vec Model Where Each Sentence Needs To Be Preserved

    See Annotator Dependencies Section in Standford Core NLP Site Below to Learn Required Preprocessing to Build a Correct Sequence of Your Data Pipeline ! *****


    NLP Preprocessing: Stemming-Lemmatization for Text Analysis:

  • Common NLP Preprocessing Task: Stemming-Lemmatization in python
  • Common NLP Preprocessing Task: Stemming-Lemmatization from Stanford NLP Group

  • See Python spaCy Libs for stemming and lemmatization below !



  • General (Short) NLP English Word List for Lexicon Based Text Analysis Applications:

    Stop word List
    Positive word List
    Negative word List


    To Get Lexicon Word (Dictionary) Lists
    WordNet -- Largest Lexical Database of English
    WordNet and Synset Download




    Basic NLP Processing for Text Analysis
    NLP 101: Tutorial for Document Vectorization Process





    Problems (Limitations) with Lexicon Based Approach with TF-IDF Ranking Score for Text Analysis:

  • Phrase (N-Gram Word) Identification
  • Can NOT Identify Relationships among Terms - Synonyms (Similar Meaning) or Antonyms (Opposite Meaning) of Terms
  • Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence - Polysemy Problem !
  • The order of Terms(Words) in a Sentence Is NOT Considered !!
  • Can NOT Identify New Terms or Changining Relationship Between New Terms and Old Terms




  • Solutions for the Probelm that TF-IDF Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence

    Context Aware Natural Language Processing (NLP) Methods:

  • POS (Part of Speech) Tagging:

  • Lecture Notes on POS (Part of Speech) Tagging (Stanford)
    English Grammar
    UPenn Treebank POS Tagging List




  • NLP NER (Named Entity) Tagging:

  • Lecture Notes on NER (Named Entity Recognizer) (Stanford)

    Software NER Tagging Parser:

    Stanford NLP NER (Named Entity) Tagging






    CoreNLP Run Test the Examples of NLP Parsers Here !



    Important Text Preprocessing Data Pipeline for Text Analysis:

    See Annotator Dependencies Section to Learn Required Preprocessing to Build a Correct Sequence of Your Data Pipeline ! *****

  • Data Pipeline to Build for Text Preprocessing (by Stanford Core NLP APIs)
  • Stanford Core NLP APIs for Preprocessing/Parsers/Taggers to Build Text Processing Data Pipeline


  • Core Stanford Parser Download for POS and NER tagging (Github repo site as well)
  • Core NLP Preprocessing/Parsers/Taggers






  • Useful Core NLP Preprocessing APIs, Parsers (Taggers/Annotators) to Download for Text Analysis

  • NLTK (Natural Language Tool Kit)



  • Tutorial for Basic NLP Processing for Text Analysis:
    Easy Tutorial for Word2Vec with CBOW and Skipgram Training Model



  • How to Set up TextBlob: Python Text Processing API
  • Google Colab Notebook: How to Set up TextBlob: Python Text Processing API
  • TextBlob: Python Text Processing API for N-Gram Identification
  • TextBlob: Another Python Text Processing API



  • SpaCy for Text Processing NLP Libs in Python :

    You can Also Use Python NLP Lib called SpaCy to build a Data Pipelining with NLP APIs for Text Analysis

  • Download and Installing spaCy
  • Introduction to spaCy101 for Text Proceesing with NLP Methods
  • spaCy: Building Data Pipleines for Text Proceesing with NLP Methods
  • Python spaCy Objects for Text Processing in Data Pipelines
  • Tutorials of spaCy for Text Proceesing with NLP Methods







  • WordNet Site (by Princeton) to downlaod NLP Software:

    WordNet by Princeton

    Download WordNet/Synnet for Sentiment Analysis on Aspect/Feature Opinion Analysis







    Semi-Supervised Learning:


    Well Known Ranking Algorithms for Vectorization of Each Document for Ranking or Labeling Ground Truth of a Training Set for a Machine Learning Classifier):

    Example of Lexicon Based Sentiment Analysis Proejct with Semi-Supervised Learning:

    2020 Presidential Election Prediction from Social Network Twitter Analysis






    Well Known Ranking Algorithms to Measure a Sentiment Score of each Token Word for Vectorization of Document for Labeling (Ground Truth) of a Training Set for Sentiment Analysis:

    Semi-Supervised Learning for Classification:

    1. Senti-WordNet for General English Texts from Page Ranking Algorithm

    Senti-WordNet for Sentiment Analysis
    Presentation of Paper: Senti-WordNet for Sentiment Analysis
    Paper: Senti-WordNet for Sentiment Analysis




    2. VADER (Valence Aware Dictionary and sEntiment Reasoner) : Social Media Text Specific Ranking Algorithm

    Ranking Algorithm: VADER for Sentiment Analysis
    Compound Score of VADER Ranking Algorithm for Sentiment Analysis
    Paper: VADER for Sentiment Analysis


    Basic NLP Processing for Text Analysis
    NLP 101: Tutorial for Document Vectorization Process

















    6-10



    Machine Learning for Data Anaytics

    Supervised Learning

    Classification with Machine Learning:


  • Machine Learning Algorithms with Probabilistic Approach:

  • Lectures:

  • Decision Tree:
  • Lecture Notes_10: Lecture Notes on Basic Classification and Decision Tree by Kumar Book

    Corrections in Lecture Notes !

    SplitINFO as Complexity Penalty

    Lecture Notes_12: Lecture Notes on Basic Classification by J. Han Book

    Lecture Notes_10_1: Lecture Notes on CART wth Case Study

    Example of Decision Tree for Fault Detection with Overfitting Problem and Post Pruning


    scikit-learn Python Implementation of Decision Tree:
    scikit-learn: Classification with Decision Tree




  • Instance Based Lazy Learners
  • Naive Bayes, KNN

  • Lecture Notes: Alternative Classification: KNN, Naive Bayes, Ensemble Methods (Kumar Book)


  • Distance Based ML Algorithm
  • K-Nearest Neighbor Algorithm

    Supporting Lecture Slides on Classifications



    scikit-learn site on Naive Bayes:
    scikit-learn: Classification with Naive Bayes




  • Ensemble Methods:

  • Lecture Note (Kumar book) on Ensemble Method with Bagging, Boosting

    Lecture Note (JHan book) on Ensemble Method with Bagging, Boosting





    Well-Known Optimized Decision Tree Based Ensemble Machine Learning Algorithms

  • Random Forest
  • XGBoost (Extreme Gradient Boost)
  • Light GBM (Gradient Boosting Machine)
  • Ensemble Method: Adaboost



  • scikit-learn Python Implementation of Random Forest:
    Ensemble Method: Random Forest (based on Bagging with CART)



    Complete List of scikit-learn Supervised Learning Algorithms











    10


    Error Estimation and Model Evaluation

    Lecture Notes_14: Lecture Notes (Combined from Both Textbooks) on Error estimation and Validation
    Corrections in Lecture Notes !
    Multi Class Model Evaluation and Comparison: Macro and Micro F1: Precisions/Recall

    Useful Tool for Model Evaluation:

    Model Evaluation and Comparison: Example of ROC with Precision and Recall
    Model Evaluation and Comparison: Macro and Micro F1: Precisions/Recall


    11-14



    Machine Learning with Numerical Approach:

  • Regression/Logistic Regression
  • Artificial Neural NetWork
  • SVM (Support Vector Machine)




  • Lectures:


    Regression as a Prediction Model:

    Lecture Note: Linear Regression
    Lecture Note: Linear Regression (Standford Tutorial)
    Introduction of Regression Models



    Basic Tutorial of Underlying Concept of Regression as Classifier:
    Regression Model






  • Logistic Regression and Classification

  • Lecture on Logistic Regression and Classification

    Short Summary of Logistic Regression

    Lecture Note: Logistic Regression (Standford Tutorial)

    Lecture on Gradient Descent Search and Regularization








  • Artificial Neural NetWork and SVM (Support Vector Machine)


  • Overview on Concept of Artificial Neural Network, Support Vector Machine (SVM) (Kumar's Textbook)


  • Artificial Neural NetWork


  • Overview of ANN Architecture at First Glance

    Lecture Notes on Neural Network to Know Better (Standford Tutorial) *************************************


    Tutorial of Step by Step Calculations with an Example on Backpropagation Algorithm in MultiLayer Neural Networks ************************

    Letcure Note on Step by Step Matrix Computation of Backpropagation with an Example ****************************

    Lecture on Basic Notation, Logistic Regression and intro to Gradient Descent and Neural Network (From DeepLearing.AI by Stanford Coursera)

    More Detail on Neural Network with Matrix Implementation of Backpropagation Algorithm (From DeepLearing.AI by Stanford Coursera)

    Standard Notations for Deep Learning



    How to Implement Forward and Back Propagation of NN for Classification ******************************************

    Implementation to Understand Calculation for Each Parameter:

    Step by Step Tutorial: Computation of Backpropagation with an example

    Sample Codes for Simple NN Implementation for Each Parameter in Python


    Matrix Implementation: ***********************************************

    Lecture: More Detail on Neural Network with Matrix Implementation of Backpropagation Algorithm (From DeepLearing.AI by Stanford Coursera)

    Tutorial and Sample Codes for Implemenation of BackPropagation of NN and Classification Using Matrix Operations in Python

    Standard Notations for Deep Learning




    Gradient Descent
    Batch Gradient Descent vs Stochatic Gradient Descent


    How to Train NN When You Don't Have a Large Training Data Set:
    Difference between iterations with batch size and epochs





    Tutorials for Classification with NN:

    scikit-learn: Classification with Neural Network


    Tutorials for Data Preprocessing Methods for NN:

    ML Algorithms -- ANN or SVM Reqire Data Preprocessing:
    Normalization for Any Numerical Attributes
    Binarization (One Hot Encoding) for Any Categorical Attributes

       Tutorial for Data Preprocessing Normalization and One Hot Encoding for ANN or SVM

    Categorical Data transformation Methods with Binarization (One Hot Encoding)

    Data Preprocessing Methods for ANN or SVM


    Tutorial for How to Implement Classification with NN in Python Implementation:

    Walk Through with Basic Example using Panda for Classification with Neural Network **********************

    Easy Tutorial on Neural Network Concept

    Walk Through with Basic Python Example in Keras Framework for Classification with Neural Network

    TensorFlow Example of Classification with Neural Network

    Keras with Tensorflow for ANN





    Easy Tutorial on Multi Layor Neural Network with Softmax Function

    Easy Tutorial on Softmax Function in NN








  • SVM (Support Vector Machine)

  • Overview on Concept of Artificial Neural Network, Support Vector Machine (SVM) (Kumar's Textbook)



    SVM Preprocessing Guides and Kernel Functions
    SVM Solutions When Not Linearly Separable

    scikit-learn SVM


    SVM Tutorial for Python Implementation:

    Understanding SVM Algorithm Deeper for Classification with Python Example
    Walk Through with Basic Example for Classification with SVM

    scikit-learn: Classification with SVM with Code Examples

    scikit-learn: Classification with SVM

    SVM Kernel Functions















    The Most Important Things to Learn for Classification:

    - Preprocessing methods to transform bigdata to accurate vector form
    - Understanding the algorithms with the differences in input parameters
    - Understanding the problems, limitations of each ML, and important factors that affect a ML model accuracy and - Solutions to fix the problems and limitations
    For example, Try multiple classifications with SVM and different kernel functions to see which kernel function generates the best model in accuracy.


    Important Problems in Classification

  • Data Preprocessing Methods for Classifiers:
  • Example of Data Preprocessing for Classification Training with Neural Network: Normalization and Binarization
    Examples of Preprocessing of Categorical Data for Classification: Appointment Cancellation Prediction
    How to Handle Unbalanced Class for Classification
    Examples of Preprocessing DateTime as Features for Classification: Appointment Cancellation Prediction
    Padas DateTime (TimeStamp) Data Type Preprocessing Functions


  • Curse of Dimensionality
  • Feature (Selection) Reduction Methods
    - PCA (Principle Compoment Analysis) -- Traditional Statistical Method
    - Correlation Analysis
    - ML Based Recursive Feature Elimination (RFE) -- Advanced ML based Method



  • Model Overfitting and Underfitting Problem



  • Class Imblance (Minority Class) Problem
  • Handling Imbalanced Data for Classification:
    Learn how to do random sampling with SMOTE for Imbalance Data for Anomaly Detection

    Synthetic Minority Oversampling Technique (SMOTE)

    Imbalanced-Learn Library
    SMOTE for Balancing Data
    SMOTE for Classification


    Handling Imbalanced Data with SMOTE in Python
    Tutorials with Examples:
    Common Sampling Methods to Handle Imbalalnced Class in Python
    Handling Imbalanced Data with SMOTE in Python



  • Small Data Set Problems
  • Random Sampling Methods
    (See at the End of the Lecture Notes on Model Evaluation)
    - Bootstrap Sampling (Sampling with Replacement)
    - Stratified Sampling
    - Stratified BootStrap Sampling


    Resampling Methods for Small Sample Problem for Classification:

    - Monte Calro Simulation
    - SVM SMOTE










    Tutorial Sites in R for Classification Steps:

    R Examples:
    Code Examples of Data Cleaning in R:
    Data Analytics (Data Mining) with R
    Multi-Class Classification with Neural Network in R







    Application Examples of Classification Task Implementation:

    Step by Step of Classification Task with Code Walk Through from Real Life Sample Project:

    Example of Application for Intrusion Detection
    Example Codes of Classification Process for Intrusion Detection
















    15-16




    Upsupervised Learning



    Clustering


    Lecture Notes_14: Lecture Notes on Clustering Analysis

    Lecture Note on K-Mean Algorithm



    scikit-learn Clustering


    Tutorial on Clustering Analysis with Mixed Data Types and Visulaization in t-SNE (t-distributed stochastic neighborhood embedding)







    Hot Spot Anlaysis for Geospatial Data Analysis

    Lecture:

    Lecture: Hot Spot Analysis in PAI and PEI Metrics to Evaluate Hot Spot/Cold Spot


    What is Hotspot Analysis and Z-score and p value based metric


    Wiki on Hot Spot Analysis



    National Institute of Justice (NIJ) Crime HotSpot Analysis Challenging:

    NIJ (National Institute of Justice) Crime Forecasting Challenge Overview



    Research Project Report on PREDICTIVE HOTSPOT MAPPING ANALYSIS (The Paper that Created PAI and PEI Metrics to Evaluate Hot Spot Analysis)

    Documents/Related Materals for NIJ Crime Forecasting Challenges:


    Documentation in Depth for Geospatial Statistics Analysis

    Data Sets:

    NIJ Challenging Data Set: Crime Data
    NIJ Challenging Data Site: NIJ Crime Location Data Set for Hot Spot Analysis
    NIJ Challenging Data Set: Crime locations in Portland


    For Those Who Want To Apply for NIJ Crime Forecasting Challenges:
    FAQs for NIJ Crime Forecasting Challenges





    GIS Data Visualization API:

  • How to Visulaize NIJ GIS Data in QGIS by Asanka Mananayaka
  • NIJ GPS data with ArcGIS

  • Useful sites for GeoSpatial data visualization map using ArcGIS API

    NIJ GPS data with ArcGIS
    Example Project to Guide How to Process Geo Spatial Data (generated from a machine) for Hot Spot Analysis From CIS 660 Project by Sarvesh Chande
    Lab2 GIS Data Processing Example of NIJ Geo Spatial Data Visulaization Using ArcGIS Map for Hot Spot Analysis
    Example of NIJ Geo Spatial Data Visulalization Using ArcGIS 3D Map for Hot Spot Analysis
    Example of GIS Data Processing in Java Script to Visualize in HTML







    Any Subject Beyond This Section Will Be Covered in CIS660 Data Mining and Machine Learning



    13







    Advanced Text Analysis with Natural Language Processing

    Semi-Supervised Learning (Self Learning) for Text Analysis:

    Context Aware NLP Methods Using Machine Learning with Semi-Supervised Learning: Word2Vec (Google)

    Problems of Learning a Sequence of Tokens in a Sentence/Phrase in the Lexicon Based Approach

    Lecture on Earlier Approach: Positioning Index for Phrase Query (Stanford)






    Natural Language Processing (NLP) with Deep Learning for Word Embeddings:

    Word2Vec:

    Context Aware NLP Methods Using Machine Learning with Semi-Supervised Learning: Word2Vec

    Lectures:

    Lecture Notes_19_1: Lecture Notes on Natural Language Processing (Stanford): Word to Vector Model (Word2Vec) in Skip Gram Model

    Tutorial on Word to Vector (Word2Vec): Skip Gram Model




    Google Site for Word2Vec Implementation and Word2Vec Papers by Google AI

    Word-2-Vec Optimization of Training Computation

    Tutorial on Optimization of Word to Vector (Word2Vec) Training



    Extended Word2Vec: Glovec
    Sites for Glove Word Vector by Stanford NLP Team

    Word2Vec Implementation Sites:

    To get an Executable Binary of word2Vec Model Implementation and Training Data sets:

    Sites for Word To Vector by Google

    Sites for Word To Vector for npmjs
    Sites for Glove Word Vector by Stanford NLP Team
    CNTK: Stanford NLP Sentiment Analysis



    Tutorial Sites for Word2Vec Implementation:

    Tutorials for Word To Vector: Skipgram Model
    Tutorials for Word To Vector with Google Tensorflow
    Tutorials for Word Embeddingd with Google Tensorflow

    Tutorials for Word To Vector and related APIs, Gensim (LDA)
    Tutorials for Bag of Words, Word2Vec and related APIs








    BERT (Bi-Directional Transformer)for Deep Learning :

    State of Art NLP AI: Google BERT Transformer
    (The Most Recent and Most Superior Word Vector Generator: Google BERT

    Tutorial: What is Google AI BERT Transformer?

    BERT from Google AI




    Tutorails:

    Tutorials on BERT

    Simple Tutorial for Sentiment Analysis with BERT


    Code Tutorial Site to Use Pretrained BERT Transformer:

    Good Tutorial Site to Use Pretrained BERT Transformer

    Good Tutorial to Use Pretrained BERT Transformer for Question Answering System

    Good Tutorial to Use Pretrained BERT Transformer for Domain Specific System

    Good Tutorial Site to Use BERT Transformer for Sentence Classification

    SpaCy Embeddings for BERT Transformer






    15


    Deep Learning:


    Basic Review For Multimedia Files:
    Basics on Color Representation: 24 bit RGB Code in 3 Color Channels (One Byte Code (2^8: 0 ~ 255) per Each Channel)
    Basics on Digitizing Images with Resolution
    Basics on Color Image Quantization
    Basics on Digitizing Sound



    Deep Learning for Image Recognition

    Lecture Notes:

    Basics Methods BEFORE Deep Leraning with CNN:

    Lecture on Basic Notation, Image File as Input with Logistic Regression and intro to Gradient Descent and Neural Network (From DeepLearing.AI by Stanford Coursera)

    Tutorial on Basic Image Recognition with Classifier



    Tutorial for Multi-Class Classification with Multinomial logistic regression for Object Identification

    Easy Tutorial on Multi Layor Neural Network with Softmax Function

    Easy Tutorial on Softmax Function in NN




    Lecture on Deep Learning with CNN *****************************************************

    Deep Learning with Convolution Neural Networks (CNN) for Classification of Image Identification (Stanford) **********************************




    Good Image Classification Tutorial sites:

    Tutorials for 10 image classification

    OpenCV Tutorial: Open Source Computer Vision in C (Python or JavaScript Wrapper Available)

    Python Deeep Learning API: Keras Tutorial

    Tensorflow with Keras Tutorial
    Recurrent Neural Network(RNN) in Tensorflow with Keras Tutorial


    Examples of Image Classification Basics to Start With:

    Tutorials for Google Tensorflow Image Recognition with CNN



    CIFAR-10 and CIFAR-100 Image data set

    Example of a Filter for CNN


    Tutorials for Google Tensorflow with CNN

    Google Analytics and AI

    Google AI Tools

    Google Xception for Image Recognition


    Useful Sites for Face Recognition:
    Open Face Research Site on Face Recognition Techniques Using Machine Learning
    Step by step Tutorials on Face Recognition Techniques Using Machine Learning
    Python with Fully Pre-configured VM for Face Recognition Data Analytics
    Simple Tutorial Site for Face Recognition With Deep Learning CNN

    Deep Learning Data Set Sites:

    Good Deep Learning Data Sets
    Open Image Data sets
    Kaggle Data sets


    16

    Presentation of Projects

    ==> Completion of Homeworks/Labs is required for obtaining a passing grade.

  • This is a tentative scale and
    it could be changed

    Letter
    Grade

    Quality Points

     


      A

    > 93%  

    A: Outstanding (student's performance is genuinely excellent)

      A-

    90% - 93%

     

      B+

    87% - 90%   

     

      B

    82% - 87%

    B: Very Good (student's performance is clearly commendable but not necessarily outstanding)

     

      B-

    80% - 82%

     

     

      C

    75% - 80%

    C: Good (student's performance meets every course requirement and is acceptable; not distinguished)
        D 65%-75% D: Below Average (student's performance fails to meet course objectives and standards)

     

      F

    <65%

    F: Failure (student's performance is unacceptable)

    ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106.

     


    Programming standards

    • Every program must include your name, CSU ID number, Class, Section Number, Hours, the words 'Homework # ...', and a short description of the assignment. For example:
       ' Name: Mark Zuckerberg  
       ' ID: 1234567            
       ' Homework #1            
       ' Description: Computing the average life of a light bulb
    • Every variable should have a meaningful name (this includes function/procedure/subprogram names).
    • Every portion of the program should be as cohesive (single purposed) as possible. This leads to a large number of small functions.
    • Every function (including the main function) should be preceded by a comment indicating its arguments and a description of the transformation it performs.
    • Non-obvious code within a function should be explained.
    • Code should not be over commented.