DSA460/CIS492/593 Data Mining and Machine Learning (4-0-4) |
||
| Course Content |
|
|
|
You can reach the class webpage from the Big Data Research Lab website as well. //a href="https://eecs.csuohio.edu/~sschung/"> Big Data Research Lab The Final Exam Info: For the Active Lecture Notes Links Only for the Final Exam, Access the following site: The rest of links are all disabled. //a href="https://eecs.csuohio.edu/~sschung/DSA460/DSA460Sp24_FinalOnlyLinks.html"> the Active Lecture Notes Links Only for the Final Exam Hand Written One Page Note (Front and Back) is Allowed to the Final. Copy and Pasted Printout of Lecture Notes Are NOT Allowed The Exam Format and Rules will be the same as the Midterm. Note that You are responsible ONLY for the Lecture Notes and the Subjects Covered in Class Question Types: 6-7 Short Answer Questions on Problem Solving with Given Small Data Sets No T/F Questions, No Multiple Choice Hand Written One Page Note (Front and Back) is Allowed to the Final. Copy and Pasted Printout of Lecture Notes Are NOT Allowed The Exam Format and Rules will be the same as the Midterm. Subjects to Focus On for the Final One question From the Midterm Questions on Classification Algorithms - KNN Advanced Classification: NN, Feedforward and Backpropagation Algorithm Calculation Clustering Algorithms - K-Mean (Tentative) Accuracy Estimation and Validation, Model Evaluation Methods Basic Deep Learning Architecture with CNN Midterm Info: Basically whichever Subjects on the Lecture Notes Covered in Class! IMPORTANT WARNING !! If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterm and the Final !! //a href="CIS460Sp24_MidtermOnly_Links.html"> Midterm Info and The Lecture Note Links Only for Midterm Exam 1 will be Tentatively on the Last Week of Feb ! Subjects to Focus on: Chapter 2, 3 on : Basic Stats of Data, Data Preprocessing Methods, Data Transformation Methods, All the Data Proximity Measures, Feature Selection Methods, Feature Correlation Measures -- Kai Square, Correlation, Covariance Basic Text Analysis Algorithms and Methods in Information Retrieval: TF-IDF Measure, Document as Term Vector See the Topics to Focus on for the Midterm (Exam 1 and Exam 2) here: //a href="CIS460Sp24_MidtermOnly_Links.html"> Midterm Links Only (All the links to the rest of the non-related lecture notes are disabled) Exam 2 will be Tentatively on the Last Week of March ! Subjects to Focus on: Classification, Machine Learning Algorithms: The Subjects Covered on Machine Learning Algorithms and Their Objective Functions Covered in Class Decision Tree, K-NN, Naive Bayes Linear Regression, Logistic Regression, ANN Feed Forward, Backpropagation Algorithms, SVM Model Evaluation Methods and Common Problems in ML Models for Classification, Data Preprocessing Methods of Each Machine Learning Algorithm for Classification Note that You are responsible ONLY for the Lecture Notes and the Subjects that Were Covered in Class Question Types: 6-7 Short Answer Questions on Problem Solving with Given Small Data Sets No T/F Questions, No Multiple Choice Hand Written One Page Note in Each Side (Only for Equations and Formulas) Is Allowed for the Midterm. 7. Jan 12, 2026: IMPORTANT NOTE !!! If You Have a 404 Error in any URL in the Class Lecture Notes under http://eecs.csuohio.edu/~sschung/, You have to replace "http" with "https" to be able to access ! 6. Jan 12, 2026: Only the registered students can access the course blackboard. If you have a problem with your blackboard access, please contact the register and CSU e-learning Tech Support to resolve the issue ! Faculty do not control your registration in the CSU Campus Systems and the course blackboard access. //a href="https://www.csuohio.edu/center-for-elearning/technical-support"> e-learning CSU Tech Support 5. Jan 12, 2026: It is Required to Attend Every Class ! Your Class Attendance Will be Checked in Each Class. There Will be Random Quizzes to Check the Attendance ! Only 5 - 10 Mins Will Be Allowed for Each Quiz. Those Who Come in Class Late After 5 Mins, They Will NOT Be Given a Quiz ! 4. Jan 12, 2026: The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class Blackboard ! Wait For the Lab Submission Links To Be Created on Blackboard with the Deadline to Submit Each Lab 3. Jan 12, 2026: TA Information: TA: Abay Suleimenov Email: a.suleimenov@vikes.csuohio.edu Office Hours: Tues, Thur 10:00AM - 12:00PM or by Appointment for a Zoom meeting Make sure to Send him email ahead to let him know you are coming. Location: Big Data Lab FH305 or ZOOM Meeting ZOOM Meeting ID: Passcode: //a href=""> If you have questions in Labs or grading your Labs, Send an email to TA to Talk During his TA Office Hours or Schedule a Zoom meeting Dr. Chung's Office Hours: Tues and Thursday 1:30PM - 3:30PM Office: FH 222 or Zoom Meeting Email: s.chung@csuohio.edu Send Me Email to Set Up a Meeting or a Zoom Meeting. Meeting ID: 859 3867 6332 //a href="https://csuohio.zoom.us/j/85938676332"> Zoom Meeting Link 2. Jan 12, 2026: The Output of each lab is your report in Doc file that shows your screen captures of your data processing steps for analytics and the results Each of your screen capture must show each step and the result returned by the server in the SAME window in your System to prove that your lab is done correctly in YOUR SYSTEM !! Lab Submission: 1) Submit your Lab in zip file including 1) your lab report in .doc and 2) all the source files, preprocessed input files, outputs on Blackboard for a timestamp as a proof. 2) If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Heading of the Front Page of Your Report in Bigger and Bold Font ! 1. Jan 12, 2026: br> Semester Schedule: //a href="https://www.csuohio.edu/registrar/academic-calendar"> See University's Official Academic Calendar for the Semester and the Final Exam Schedules Final Exam Schedule: Tue May 5 4:00p-6:00p Exam Info: Basically whichever Subjects on the Lecture Notes Covered in Class! IMPORTANT WARNING !! If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterms and the Final !! //a href="CIS460Sp24_MidtermOnly_Links.html"> Midterm Info and The Lecture Note Links Only for the Midterms There Will Be 3 Exams (15% Each) Subjects to Focus on: Chapter 2, 3 on : Basic Stats of Data, Data Preprocessing Methods, Data Transformation Methods, All the Data Proximity Measures, Feature Selection Methods, Feature Correlation Measures -- Kai Square, Correlation, Covariance Basic Text Analysis Algorithms and Methods in Information Retrieval: TF-IDF Measure, Document as Term Vector See the Topics to Focus on for the Midterm (Exam 1 and Exam 2) here: //a href="CIS460Sp24_MidtermOnly_Links.html"> Midterm Links Only (All the links to the rest of the non-related lecture notes are disabled) Subjects to Focus on: Classification, Machine Learning Algorithms: The Subjects Covered on Machine Learning Algorithms and Their Objective Functions Covered in Class Decision Tree, K-NN, Naive Bayes Linear Regression, Logistic Regression, ANN Feed Forward, Backpropagation Algorithms, SVM Model Evaluation Methods and Common Problems in ML Models for Classification, Data Preprocessing Methods of Each Machine Learning Algorithm for Classification Note that You are responsible ONLY for the Lecture Notes and the Subjects that Were Covered in Class Question Types: 6-7 Short Answer Questions on Problem Solving with Given Small Data Sets No T/F Questions, No Multiple Choice Hand Written One Page Note in Each Side (Only for Equations and Formula) Is Allowed for the Midterm. For the Active Lecture Notes Links Only for the Final Exam, Access the following site: The rest of links are all disabled. //a href="https://eecs.csuohio.edu/~sschung/DSA460/DSA460Sp24_FinalOnlyLinks.html"> the Active Lecture Notes Links Only for the Final Exam Hand Written One Page Note (Front and Back) is Allowed to the Final. Copy and Pasted Printout of Lecture Notes Are NOT Allowed The Exam Format and Rules will be the same as the Midterm. Subjects to Focus On for the Final One question From the Midterm Questions on Classification Algorithms - KNN Advanced Classification: NN, Feedforward and Backpropagation Algorithm Calculation, Ensemble Methods Clustering Algorithms - K-Mean (Tentative) Accuracy Estimation and Validation, Model Evaluation Methods Basic Deep Learning Architecture with CNN Important Note for the Schedule of the Midterm Exams and Final Exam There is always possibility that 2-3 course exams would be clustered in the same day during the Midterm Weeks or Final Exam week. |
DSA 460 Data Mining and Machine Learning is a Core Required Course of the Data Science Degree (BSDS)
DSA 460 Is Also Offered as CIS492/593 Data Mining and Machine Learning for the CS Majored Students as an Advanced Elective Course
|
Labs: First Week Lab0: Choose your System/Tool/Platform to Set Up and Get Used to: See the Lab Assignment 0 Section Below to See the Set Up Guide for Python Data Science Platform in Step by Step Basic Python Tutorials : Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Python Data Science Platforms: There Are Mainly Three Ways to Set up All the Neccessary Data Science Software/Library ! See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !! //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) //a href="https://www.dataquest.io/blog/jupyter-notebook-tutorial/"> Basic Guide for Python tools for Data Science: jupyter-notebook //a href="https://www.dataquest.io/blog/advanced-jupyter-notebooks-tutorial/"> More Basic Guide for Python tools for Data Science: jupyter-notebook • Anaconda //a href="https://www.anaconda.com/open-source"> Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms //a href="https://docs.anaconda.com/anaconda/navigator/tutorials/"> Anaconda Tutorials • Python Anaconda Tutorial Sites //a href="https://data-flair.training/blogs/python-anaconda-tutorial/"> Anaconda Tutorial Site //a href="https://www.edureka.co/blog/python-anaconda-tutorial/"> Anaconda Tutorial Site • PyTorch //a href="https://pytorch.org/"> PyTorch Site (It can be integrated from Anaconda as well) Google Colab: • Python Scikit Learn for Common Data Science Tasks Text Preprocessing (Natural Language Processing) Library in Python SpaCy: //a href="https://spacy.io/api"> Liquistic Modules in Python SpaCy //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing • Python sklearn.cluster //a href="https://scikit-learn.org/stable/modules/clustering.html"> Python Sklearn Clustering There is Another Data Science Platform in R if You Choose to Learn (SAS Data Miner or MS Data Tool Have an Integrated R Platform) Basic R Tutorials : //a href="IndependentStudyCIS611Final Report.pdf"> Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study by Nick White (Now in FaceBook) Useful Machine Learning Tutorial Sites: //a href="https://keras.io/guides/"> Keras for Image Processing/Text Processing with Deep Learning //a href="https://colab.research.google.com/"> Google Colab for Fast Machine Learning Execution in GPU Good Data Science and AI Platforms: For Your Own Study Lab Submission Instructions: The Output of each lab is your report in Doc file that shows your screen captures of your system/tool/platform setting/configurations, your data processing steps for analytics and the results Each of your screen capture must show each step and the result returned by the server in the SAME window in your System to prove that your lab is done correctly in YOUR SYSTEM !! 1. Submit (on Blackboard) your Lab in a zip file including 1) your lab report in .doc and 2) all the source/scipt files, preprocessed input files, outputs on Blackboard for a timestamp as a proof. 2. If you did an Extra Credit Lab, Make a Note on the COVER of Your Lab Clearly ! The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font ! Always Follow the Deadline of Each Lab Assigned on the Class Blackboard. The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester. Please Identify Your Course When You Ask Me in Email ! 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof. 2. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System Name to Prove That Your Lab Was Done by YOU on Your Computer. 3. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font ! 4. Always Follow the Deadline of Each Lab Assigned on the Class Blackboard !! The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester. Lab Assignment 0: Set Up Your Python Data Science Platform for Labs and Final Projects by the End of the First Weekend ! There Are Mainly Three Ways to Set up All the Necessary Data Science Software/Library ! See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !! (Not for This Semester !) PreLab 0 : Preliminary Analysis of the Computer Game Sales Data (Not for this Semester !) With the Given Table - Global Video Game Sales data (below) obtained from the Data Warehouse of a Game Retailor Company, Analyze the Game Sales Data to Identify Any Trend in the Total Sales of the Game During Last 10 Years. You need to find the following: Which Computer Game Genres, Platforms Will Be the Best to Focus to Invest/Develop for Next Year (in this data set next year is 2017) Globally and For Each Region. To be able to predict, you need to analyze sales data to provide any evidence of your prediction: Hint: 1. See total sales by Genres, Platforms respectively per each year during last 10 years. 2. Save each analyzed result into a CSV file (or another excel sheet) for further processing or any visualization tool or Excel to Visualize in Graph to Compare total sales of each Genres, Platform during last 10 years to find any useful intelligence to Predict 3. Report the facts found from your visualization of each analysis result and discuss why you predict that the specific Genres and Plaform are best to invest in for the next year. You may use Any Tool of Your Choice for analysis and visulization -- Even Excel should be enough for this task to analyze and visualize the analyzed results. Hint: One easy way to create analyzed data is that you can do all the calculation per each year and each genre , (each year and each platform as well). To create a graph, you need each aggregated total per each year and each genre, (each year and each platform as well), saved in one file (this can be done by storing the SQL result in a table), then saved the table as an excel or csv file for a graph. Analyze and Visualize the Computer Game Sales Data of a Company below to Predict which Genre and Platform of Computer Games Will Be the Best Sellers Next Year. The Predicted Genre and Platform Can Be Provided as Decision Supporting Evidences for the Company to Make a Correct Decision to Invest for Next Year (in this data set, next year is 2017) Globally as well as Each Region. You may use any tool of your choice -- Even Excel should be enough for this task to analyze and visualize the analyzed results. Although the database skill is NOT required for this course, it will be very useful for handling big data sets. For Those Who want to Use a Database Server for Initial Data Handling Installation Guides for MySql Server: //a href="https://www.mysql.com/downloads/"> MySql Download //a href="https://dev.mysql.com/doc/refman/5.7/en/installing.html"> MySql Download and Installation //a href="https://dev.mysql.com/doc/refman/8.2/en/tutorial.html"> MySql Tutorial //a href="https://dev.mysql.com/doc/refman/5.7/en/creating-database.html"> How to Create MySQL Database Microsoft SQL Server: //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> See the Lab Sections of CIS 430/530 for Microsoft SQL Server Installation Guides Useful SQL Review to Get to Know Your Data Features (Columns) REQUIRED LAB ASSIGNMENTS: The Submission Link and the Deadline of Each Lab Will Be Posted on the Class Blackboard ! Lab Assignment 1: Lab1_1 - Part 1 and Part 2 on Preprocessing and Transformation: Due By the End of the Third Week Lab1_2 - Part 3 and Part 4 on Similarity Measure and Building Similarity Matrix Due By the End of the Fourth Week Input Data File for Lab1: For Part 1 and 2: For each selected feature, identify ALL the required data preprocessing methods like Normalization, Discretization, Binarization, and more based on the feature data properties. They must be done along with other preprocessing methods. Examples of Lab1 Report of Data Preprocessing and Data Similarity Measures //a href="Lab1_OutputExample_1.pdf"> Example1 of Lab1 Output: Part 1 //a href="Lab1_OutputExample_Part2.pdf"> Example of Lab1 Output for the Final Transformed Features Note that these Examples Here May NOT neccessarily all correct ! They just show How Lab1 can be done as an example. Do NOT blindly follow ! For example, EnglishEducation should be Transformed as Categorical or Ordinal? If it is transformed as Categorical, One-hot-encoding is a correct transformation. However, if it is Ordinal, then one hot encoding is not a correct transformation for this column. For Part 3: NOTE !!! You Are NOT Supposed to Use any Python Libs to Calculate Similarity Measures in Part 3 of Lab 1. You have to write Scripts/Programs to Compute each Similarity Measure. If you use the built-in lib to build a Similarity Measures, which is one line of code, you will get 0 for the part. Suggested Platforms to Use for Lab1: Anaconda //a href="https://www.anaconda.com/open-source"> Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms //a href="https://docs.anaconda.com/anaconda/navigator/tutorials/"> Anaconda Tutorials FAQs for Lab1: Q: For Part 1 of Lab 1, is it acceptable to complete the data processing using Excel? I wanted to confirm whether submitting the work in Excel is appropriate, or if you would prefer the responses to be documented in a Word file instead. A: It was instructed to set up to use one of the Data Science platforms introduced in class as Lab0 and create a Lab report in doc file with the screenshots showing the intermediate results. Q: Additionally, for Lab 2, should we use the dataset generated from our random sampling, or the dataset that was provided in the lab materials as the input? A: The Lab1 specification clearly mentioned at the beginning that use a full data set for the tasks in Part 2 as below: Part 2: Data Preprocessing and Transformation Use all the data rows (~= 18000 rows) with the selected features as an input file to apply the tasks below, do not perform each task on the smaller data set that you got from your random sampling. Q: In Part 2 of Lab 1, it says we need to replace the Null values. What are we replacing them with? Also, later in part 2, it says that we need to replace the outliers. I am also confused what we are replacing the outliers with. A: The Null value replacement methods to determine the value to replace Nulls were covered in the class lecture notes at: //a href="https://eecs.csuohio.edu/~sschung/DSA460/chapter_3_1_DataPreprocessingCleaningOnlyNull_Outliers.pdf"> Null Replacement Methods Depending on the feature column type, for example, Mean, Median, Min, or Weighted Mean would be a representative value to replace Null with. For an Outlier replacement, 1.5IQR to calculate the upper extreme and low extreme to identify outliers and replace them with the lower extreme, upper extreme in each side as explained at: //a href="https://online.stat.psu.edu/stat200/lesson/3/3.2"> 1.5IQR Outlier Replacement Method Lab Assingment 2: Note That Lab2 Requires Stemming/Lemmarizaion and Bi-Gram and Tri-Gram Handling You can use any avaialble API for Preprocessing Some Trouble Shooting Tips: For Lab 2, it requires the python library 'clean-text' The full command to install this dependency is: pip install clean-text If you do wish to rerun it and make sure that it works, you would have to uninstall the old one first as they use the same module name. Having both installed at the same time will result in Python using the incorrect one. NOTE !!! You Are Not Supposed to Use any Python Libs to Calculate a Cosine Similarity for Lab 2. You have to write a script to compute and build a Similarity Matrix. If you use the built-in lib to build a Cosine Similarity, you will get 0 for Labs. Lab Assingment 3 on Classifier IMPORTANT NOTES for LAB3: 1. Use the Best Selected Features Given Below instead of Your Own Selected Feature Set in Lab1 : IMPORTANT !!: In your Lab report, Make sure to Show the final transformed training set file and all the attribute values for the first two objects in your training set that was used for your classifiers. 2. 4 Classifiers for Lab3 as below: 3. Grading is Based on Your Best Accuracy with the best input parameters you identified from Your Experiments Tutorial with Codes Examples for Implementation of Classification with ML Algorithms //a href="DSA460_CIS492_593_ClassificationDTExample.pdf">Example of Classification with DT //a href="DSDA460_CIS492_593_ClassificationDTExample2.pdf">Example 2 of Classification with DT //a href="DSA460_CIS492_593_ClassificationANN_SVM_LabExample3.pdf">Example of Classification with ANN and SVM Lab Assingment 4 on ANN: Lab 4_1 on Calculation of ANN Backpropagation: Lab4_1 - ANN Calculation in Pencil and Paper — Required for All Lab4_2 - Classification with ANN - Required for All for All Extra Credit Lab4_3 (100 %): Lab4_3 - Implementation of ANN Algorithm in Feedforward and Backpropagation in Matrix Operations in Linear Algebra either in Python Numpy or MatLab for Classification Lab 4_2 on Classification with ANN and SVM: Use the Same Data Set Given for Lab3 for model comparison Requirements for Lab4_2: 1. Experiment Your Classification with two Numerical Approach Based MLs: ANN and SVM 2. For Each ML, Find the Best Hyperparameters (input parameters) to Find the Best Performing model With ANN, Repeat training to find the optimal hyperparameters (Learning rate, the Numer of Hidden Layer) for the best fit(model) for the given data With SVM, Repeat training to find the optimal hyperparameters (Kernel Function) for the best fit(model) for the given data Data Set for Lab4_2 Extra Credit Lab4_3: Data Set for Lab4_3: Patient Data Set for Breast Cancer Prediction for Binary Classification to Predict whether the Patient has a Breast Cancer or not See the Lecture Note Section for Implementation of ANN for this Implement Feedforward and Back Propagation of NN for Classification //a href="https://mattmazur.com/2015/03/17/a-step-by-step-backpropagation-example/">Step by Step Tutorial of Backpropagation with an example //a href="TC8_NN_BackPropagation.pdf"> Lecture Note on Step by Step Matrix Computation of Backpropagation with an Example ************************ //a href="C1_W3_Lecture_NN_BackPropagation.pdf">More Detail on Neural Network with Matrix Implementation of Backpropagation Algorithm (From DeepLearing.AI by Stanford Coursera) Matrix Implementaion: //a href="Standard notations for Deep Learning.pdf">Standard Notations for Deep Learning //a href="https://github.com/Kulbear/deep-learning-coursera/blob/master/Neural%20Networks%20and%20Deep%20Learning/Planar%20data%20classification%20with%20one%20hidden%20layer.ipynb"> Sample Codes for Implemenation of Forward and Back Propagation of NN Using Matrix Operations in Python Lab Assingment 5 on Clustering: Some Suggested Platforms for Clustering //a href="https://scikit-learn.org/stable/modules/clustering.html"> scikit-learn Clustering Although Clustering in general should work on data as multidimensional vectors, For NIJ Data For Lab4, you can work on NIJ Challeneges to identify Hot Spots for Crime Location - GPS data and Crime Category instead of handling data as multidimensional object for Clustering algorithms Cluster on the Crime Loactions (X and Y coordinates - transform them to GPS coordinates) with K-Mean and DBCSAN and Cluster the Crime locations for each Crime Category but Add PAI or PEI Analysis for Hot Spot Analysis with Changing Parameters National Institute of Justice (NIJ) Crime HotSpot Analysis Challenging: //a href="https://nij.ojp.gov/funding/real-time-crime-forecasting-challenge-posting"> NIJ (National Institute of Justice) Crime Forecasting Challenge: data description and overview |
|
Project: You Can Choose Your Own Data Set to Extend Lab3 to Design Your own Classification for Final Group Project ! Data Set Repository for Classification: //a href="https://www.kaggle.com/datasets"> Kaggle Data Set Repository //a href="http://archive.ics.uci.edu/ml/datasets.php"> UCI Data Set Repository for Classification //a href="https://healthdata.gov/search/type/dataset"> Health Data Set Repository for Classification //a href="https://dev.socrata.com/"> Socrata Open Data API Important Dates (Tentative) for Final Group Project (The Exact Deadlines will be Posted on the Blackboard) Group Project Proposal Due by April 3rd ! Group Project Status Report Due by April 17 ! Group Project Presentation Either on April 28 or 30 ! Final Group Project Report Due By Friday May 1 ! Submit Your Group Proposal (in Minimum 3 Pages) on BlackBoard Group Proposal Should Include: 1. Data Description, Data Size, Data Collection Plan (if needed), 2. Systems/Tools to Use, 3. Data Preprocessing Methods, 4. Data Analytic Goal with Plan of Evaluation (Design of Your Experiment) in Detail Your Project Status Report Should Show the Following Tasks Done: 1. Platform Setting/Configuration Procedure if it is new, 2. Your Data Contents, Selected Feature Description, 3. Data Preprocessing Steps and the intermediate Outputs Your Project Presentation Should Have Contents on: 1. Data Description, Data Size, Data Collection Method(if needed), 2. Goal of Your Project for Data Analytics or Data Mining Application (AI) 2. Systems/Tools to Used, 3. Feature Selection Method and The Final Features Selected 4. Data Preprocessing Methods and Intermediate Results, Final Training Set and Test Set Description 5. Data Analytic Design (of Your Experiment) in Detail Show the Source Codes for Each Below: 5_1 Preprocessing and Transformation Methods Done 5_2 Size of Training and Test Set 5_3 How Many Classes in Your Classification and Class Distribution in Size 5_4 How Training Is Done 5_5 Model Evaluation Metrics Used 5_6 Results in Accuracy and Comparisons 6. Challenges/Problems Encountered - How to Resolve Them 7. Evaluation Results and Visualization of the Results For the Class Size with < 25 Students, 1-2 Person Group Is Allowed For the Large Class with > 35 Students, Any Group Size in 1 - 4 Person Group Are Allowed Note that Large Groups (3-4 Person Group) Shoud Complete a Bigger Project ! One Submission Per Group Is Required ! EACH Memeber Name and ID SHOULD BE LISTED in the COVER ! List Your First and Last Name Only ! Exactly as Appeared in the CSU CampusNet. DO NOT Use Your Middle Name. If We Can't Find You by Your Name appeared on Your Project Report and Presentation Schedule. Your Project will be Considered as Missing with 0. Task 4: Final Group Project Report Submission Instructions (By the End of Friday of Your Presentation Week): Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week ! Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx ! If your data file is too big to upload, Submit your zip file with your Data file on your ONE drive and Send email to me and TA attached the data file on One Drive. Submit a Zip file that includes: 1. All of your presentation slides (both in .ppt and .pdf) and 2. Your Group Final Project Report (in doc) that shows with: 1.Platform/System Set up Procedures/Instructions, 2. Evaluation Results and Visualization of the Results 3. Executions Steps, all the source codes/scripts, all the intermediate outputs, and final output files 4. Include the Problems/Error Encountered and Your Resolutions in Your Report One Submission Per Group Required. Submit a Zip File that Includes All the required Source Files, Input, Output Files, and Final Report (in Doc file) and Presentation Slides (in pptx). Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files. The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes from the Web. Good Data Sets for Classification: //a href="https://healthdata.gov/stories/s/nqx6-g6vz"> Health Data Sets for Challenges **************** //a href="https://data.gov/"> US Government Data Sets **************** More Project List and Data Sets: Sentiment Analysis of Online Reviews/Social Media data: Senior Design Projects on Big Data and AI: 2022 - 2024: Big Data and AI Projects: 2020 - 2021: Big Data and AI Projects: 2019 - 2017: Big Data and Data Science Projects: For the Graduate Courses CIS593 and EEC525 **************************** More Advanced Topics for Final Project Implement Your Own (Deep) Neural Network Architecture Using Google Tensorflow or MS CNTK or any of your choice for Text Analysis Task or Image Processing: Extra Credit ! If you want to work on Neural Network, then download source codes of skipgram (the first paper) from one of those sites below to learn how to build Skipgarm with NN then implement paragraph vector as document vector in the second paper. Look at the first research project for guide for this project at: Then, let's read the next two papers to understand the source codes for you to download and do the experiment with real data set. Then check these two sites where you can download all the source codes to start an experiment with real data sets. Data set to train to generate word2Vec and Paragraph Vector: You can choose any webpage set or papers as data but it should be at least 200,000 documents to train There is a preprocessed wiki page data set is available in the Stanford NLP site that can be used as your training data set. You can make this as your project if you want. The Most Recent and Most Superior Word Vector: BERT -- See Advanced Text Analysis Section in Class Lecture Notes for more on BERT //a href="https://ai.googleblog.com/2018/11/open-sourcing-bert-state-of-art-pre.html"> BERT from Google AI //a href="http://jalammar.github.io/illustrated-bert/"> Tutorials on BERT from Google Good Tutorial Sites to Learn How to Use Pretrained BERT: //a href="https://spacy.io/usage/embeddings-transformers"> SpaCy Embeddings for BERT Transformer //a href="https://colab.research.google.com/drive/1yFphU6PW9Uo6lmDly_ud9a6c4RCYlwdX"> Good Tutorial Site to Use Pretrained BERT Transformer //a href="https://www.youtube.com/watch?v=l8ZYCvgGu0o&ab_channel=ChrisMcCormickAI"> Good Tutorial to Use Pretrained BERT Transformer for Question Answering System //a href="https://colab.research.google.com/drive/19loLGUDjxGKy4ulZJ1m3hALq2ozNyEGe"> Good Tutorial to Use Pretrained BERT Transformer for Domain Specific System //a href="https://colab.research.google.com/drive/1pTuQhug6Dhl9XalKB0zUGf4FIdYFlpcX"> Good Tutorial Site to Use BERT Transformer for Sentence Classification //a href="PassageRerankingBert2019.pdf"> Research Paper: Passage Reranking Using BERT from Google NSL-KDD data set for AnomalyDetection: //a href="https://www.kaggle.com/hassan06/nslkdd"> Kaggle: KDD data set for AnomalyDetection //a href="https://www.unb.ca/cic/datasets/nsl.html"> Data set Related Research site on AnomalyDetection For NIJ Challenging: Hot spot Analysis: See Clustering Section for more detailed information See More Info for NIJ Challegeing in the Clustering Section Download New data sets and Submission files of Winning Teams below to Learn from the NIJ (National Institute of Justice) Challenege site (Scroll down to see the links) Find Out How to Measure Hot Spots in score type PAI and PEI* for Every Category, Crime Type, time Frame. GIS Data Visualization API: Text Analysis Using Classification //a href="ISISTwitterVishnuSantosh.pdf">Text Mining on Twitter Data on ISIS terrorists group and the fan groups of ISIS Santosh Tankala, Vishnu Vishnuteja Thummanapelli, and Akhi Reddy Laxmanagari //a href="ISISTwitterVishnuSantosh.pdf">Tutorial: How To get Twitter Data Another Projects: Sentiment Analysis of Online Reviews: |
| Class | Chapter / Topic / Specific Objectives / Activities |
| 1 |
Introduction to Big Data Analytics: |
| 1-4 |
|
| 4-5 |
|
| 6-10 |
|
| 10 |
|
| 11-14 |
|
| 15-16 |
|
| 13 |
|
| 15 |
|
| 16 |
//a href="CIS660_BigDataProjectList.pdf">Presentation of Projects |
==> Completion of Homeworks/Labs is required for obtaining a passing grade.
| This
is a tentative scale and |
Letter |
Quality Points |
|
||
| A |
> 93% |
A: Outstanding (student's performance is genuinely excellent) | |||
| A- |
90% - 93% |
||||
| B+ |
87% - 90% |
||||
| B |
82% - 87% |
B: Very Good (student's performance is clearly commendable but not necessarily outstanding) | |||
|
|
B- |
80% - 82% |
|||
|
|
C |
75% - 80% |
C: Good (student's performance meets every course requirement and is acceptable; not distinguished) | ||
| D | 65%-75% | D: Below Average (student's performance fails to meet course objectives and standards) | |||
|
|
F |
<65% |
F: Failure (student's performance is unacceptable) | ||
|
ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106. |
Programming standards
|
|
|