credit risk analysis

Cancelado Publicado hace 6 años Pagado a la entrega
Cancelado Pagado a la entrega

1. Use the data in the contents tab labeled Prosper.csv. Create a good ‘risk’ model.

a. For the model building use SAS eMiner.

b. Greater thoroughness in the descriptive statistics, the writeup/documentation, the formatting of the report, and in how you came to your conclusions.

2. Descriptive stats. For each variable in the dataset summarize it (max,min,mean,etc) and make a histogram. Label and format appropriately. Do the same for any variables you create. Comment on anomalies in your data and what you did to address them. Defend your reasoning. It might be easiest to just make one page per variable. You should be very thorough in this area. There is code and some of the midterm papers and presentations posted in d2l for you to learn from. Whether you use R or SAS for the descriptive statistics is unimportant however it would seem to me that R would be best given you have worked in that environment more at this point.

3. Create a logistic regression model. In your attempt you should justify why you used the variables you did and for each rejected variable explain why you rejected it. Many groups on the midterm did a little hand waving here which is fine but now that you have gone through it you need to be more thorough and specify each variable. The variables you use in the final model should be binned and transformed to WOE. This is a change from the midterm. On the midterm there were several variables that had ‘goofy’ values. There were 999’s. There were odd values of outliers. There were categorical variables that were too many categories. A WOE transformation can ‘fix’ all of this. You can do the WOE transformation in R if you like or in eMiner the ‘interactive grouping’ node is specially made just for this task. If you use the interactive grouping node DO NOT simply accept the base grouping SAS throws at you. Go in and adjust the bins/groups in a logical fashion. Use the WOE_variables as the inputs to the model rather than the original variables.

4. For modeling you should sample the data into two groups. Typically I use 60/40 but you can use 70/30 or 50/50 if you like.

5. In addition to the above logistic regression you should attempt a more advanced model in eMiner and report its results in comparision. This might be a form of Neural Net, LARS, Dmine regression, etc. There are a number of them in eMiner. This is to show how different models can be applied in a straightforward manner once the problem has been set up correctly. For the ‘other’ model it is preferable you not use the WOE variables since those are special purpose things for logistic regression. It isn’t invalid to use WOE but it is more interesting to compare without the WOE.

6. Submit your report and any code you made. Neatness, organization, style, etc will be part of the consideration. The two nodes that can help you here are the ‘model comparison’ node which makes nice output and ROC curves and tables and such. The second node is the ‘reporter’ node which outputs a complete log of the process. In the reporter node when you are selecting the parameters you can make your life easier by telling it to create ‘rtf’ which is a ‘rich text file’ which opens in Word. Additionally you can alter the quality/type of output as shown in class.

7. We will ‘present’ during the final exam period. I will randomly draw names to present. You should prepare and submit a presentation that is not longer than 15 minutes. In other words we’ll have enough time to do several.

Final deliverables then are:

a. Code file as appropriate (or just put it in an appendix of the report).

b. Report.

c. PowerPoint presentation.

Lenguaje de Programación R SAS

Nº del proyecto: #13830753

Sobre el proyecto

3 propuestas Proyecto remoto Activo hace 6 años

3 freelancers están ofertando un promedio de $97 por este trabajo

ashitjha7

20+ years industry work experience in the area of IT, Finance & Banking. Also, I have 4+ years of experience in Economics, Business Analytics and Advanced statistics projects using software such as Python, R, SPSS, Min Más

$160 USD en 4 días
(1 comentario)
2.6
ExpertzWorld

“Time. It’s what we writers fight for. Without it, we have no hope of bringing our written creations to life. We need time to study, time to read, time to ponder, time to dream, and of course – time to write.” I am Más

$30 USD en 1 día
(0 comentarios)
0.0
techfinity3

DDear Prospect Hiring Manager. Thank you for giving me a chance to bid on your project. i am a serious bidder here and i have already worked on a similar project before and can deliver as u have mentioned I have Más

$208 USD en 6 días
(0 comentarios)
0.0