Showing posts with label Automated Valuation. Show all posts
Showing posts with label Automated Valuation. Show all posts

Monday, September 28, 2020

How to Analyze and Present a Complex Dataset – in 60 Minutes

        For New Graduates/Analysts

We talked about analyzing and presenting a large and complex dataset in 30 minutes in the prior blog post. Would one handle it differently if one had 60 minutes? Here is one approach one might like to consider:

1. While starting out, many young folks tend to underestimate themselves. The very fact that one has been tasked with this critical presentation speaks volumes, so one must learn to take full advantage of this visibility in narrowing the (internal) competition down. These meetings are often frequented by other department heads and high-level client representatives, leading to significant loss of time in unrelated (business) discussions. The best way to prepare for such contingencies is to split the presentation into a two-phase solution where phase-1 leads seamlessly to phase-2. 

2. In a business environment, it's never a good idea to start with a complicated stat/econ model; instead, one must start a bit slow but use one's analytical acumen and presentation skill to gradually force people to converge on the same page, retaining maximum control over the presentation in terms of both time and theme). Therefore, the phase-1 solution should be the same as the full 30-minute solution we detailed in a prior blog post (including the sub-market analysis). Even if the meeting leads to unrelated business chit-chat, off and on, the presenter will still be able to squeeze in the phase-1 solution, thus offering at least a baseline solution. Alternatively, if one has an all-encompassing solution, one could end up offering virtually nothing. 

3. Now that the phase-1 presentation, establishing a meaningful baseline is over, one should be ready to transition to the higher-up phase-2 solution. In other words, it's time to show off one's modeling knowledge. The phase-1 presentation comprised a baseline Champ-Challenger analysis, where the Champ was the Monthly Median Sale Price, and the Challenger was the Monthly Median SP/SF. The presenter used the "Median" to avoid having to clean up the dataset for significant outliers. Here is the caveat of sales analysis though: Sales, individually, are mostly judgment calls; for example, someone bent on buying a pink house would overpay; an investor would underpay by luring a seller with a cash offer, etc. In the middle (middle 68% of the bell curve), the so-called informed buyers would use five comps, usually hand-picked by the salespeople, to value the subjects – not an exact science either.  

4. Now, let's envision where the presenter would be at this stage – 30 minutes on hand and brimming with confidence. But it's not enough time to try to develop and present an accurate multi-stage, multi-cycle AVM. So, it's good to settle for a straight-forward regression-based modeling solution, allowing time for a few new slides to the original presentation. Ideally, the model should be built as one log equation with a limited number of variables (though covering all three major categories). The variables one might like to choose are: Living Area, Age, Bldg Style, Grade, Condition, and School/Assessing District, avoiding the 2nd tier variables (e.g., Garage SF, View, Site Elevation, etc.).

5. One should use Time Adjusted Sale Price (ASP) as the dependent variable in the Regression model, explaining the connection between the presentations (meaning phase-1 and 2) so the audience (including the big bosses like the SVP, EVP, etc.) understands that the two phases are not mutually exclusive, rather one is the stepping stone to the other. At this point, the presenter could face this question "Why did you split it up into two?" The answer must be short and truthful: "It's a time-based contingency plan."

6. At this point, the presenter must keep the regression output handy without inserting it into the main presentation, though, considering it is a log model (the audience may not relate to the log parameter estimates). If the issue comes up, the presenter should talk about the three critical aspects of the model: (a) the variable selection (how all of the three categories were represented), (b) the most vital variables as judged by the model (walking down on the t-stat and p-value), and (c) overall accuracy of the model (zeroing on the primary stats like r-squared, f-statistics, confidence, etc.).   

7. The presenter must explain the model results in three simple steps: (a) Value Step: ASPs vs. Regression values, showing the entire percentile curve, 1st to 99th percentile rather than the median values only, and also pointing out the inherent smoothness of the Regression values vis-a-vis the ASPs; (b) Regression Step: How some arms-length sales could be somewhat irrational on both ends of the curve (<=5th and >=95th) and why the standard deviation of the Regression values was so much lower than ASP'; and (c) Ratio Step: Stats on the Regression Ratio (Regression Value to ASP) as it's easier to explain the Regression Ratios than the natural numbers so spending more time on the ratios would make the presentation more effective.   

8. The presenter should explain the outlier ranges -- the ratios below the 5th and above the 95th percentile, or below 70 and above 143. Considering this is the outlier-free output, it's good to display Std Dev, COV, COD, etc. The outlier-free stats would be significantly better than the prior (with outliers) ones. Another common outlier question is: "Why no waterfront in your model?" The answer is simple: Generally, waterfront parcels comprise less than 5% of the population, hence challenging to test representativeness. (In an actual AVM, if sold waterfront parcels properly represent the waterfront population, it could be tried in the model, as long as it clears the multi-collinearity test as well).  

9. Last but least, one must be prepared to face an obvious question: "What is the point of developing this model?" Here is the answer: "A sale price is more than a handful of top-line comps. It comprises an array of important variables like size, age, land, building characteristics, fixed and micro-locations, etc. so only a multivariate model can do justice to sell prices by properly capturing and representing all of these variables. The output from this Regression model is the statistically significant market replica of the sales population. Moreover, this model can be applied to the unsold population to generate significant market values. Simply put, this Regression model is an econometric market solution. Granted, the unsold population could be comp'd, but that's a very time-consuming and subjective process."

-Sid Som
homequant@gmail.com

How to Analyze and Present a Complex Dataset – in 30 Minutes

For New Graduates/Analysts

Often, with minimal time on hand – say 30 minutes – to summarize and present a relatively large and complex home sales dataset, comprising 18 months of data, with 30K rows, and ten variables, here is one approach worth considering:

1. Given the limited time, instead of trying to crunch the data in a spreadsheet, it's better to one's your favorite statistical software like SAS, SPSS, etc. What SAS will do in four short statements (Proc Means, Var, Class, and Output), and a matter of minutes, will need much longer to accomplish the same in spreadsheets. When one is starting out, it's good to take full advantage of these types of highly visible, often gratifying challenges to narrow the potential competition down.

2. It's good to have a realistic game plan. Instead of shooting for an array of parameters, it's better to start with the most significant one, i.e., Monthly Median Sale Price (and the normalized Sale Price per SF). Since the median is not prone to outliers, the dataset doesn't have to be edited for outliers, saving a significant amount of time.  

3. Now that the monthly median prices are there, one should be ready to create graphs for the presentation. While one graph depicting both prices (Y1 and Y2) against months (X-axis) may be created, it's prudent to keep them separated for ease of presentation. 

4. Since basic graphing is more straightforward in Excel (in fairness to the remaining time), it's better to transfer the output from SAS to Excel, ensuring that the graphs are adequately annotated and dressed up with the axis titles, legends, gridlines, etc. One must also remember that just doing things the right is not good enough, one must learn to present things elegantly as well. 

5. Since so much of the data have been summarized and rolled up behind one or two graphs, one must make sure they not only tell the overall story but also convey enough business intelligence to make the presentation look like a well-thought-out business solution in front of the attending EVP, SVP, etc. In the presence of clients, it enhances the bosses' image as well. So, it's smart to add trendlines alongside the data trend, selecting the primary trendline by eyeballing the data trend (linear, logarithmic, polynomial, etc.). Adding a 2 to 3-month moving average (depending on the time series) trendline to iron out any monthly aberrations could enhance the presentation.

6. It's also smart to keep the reporting verbiage clear and concise, explaining the makeup of the dataset, methodology including monthly medians, and how the normalized prices add value and help validate the primary. It's also important to explain the use of the trendline and its statistical significance and the other statistical measures like r-squared, slopes, etc. one might display on the graphs (avoiding the printing of equations on the graphs). 

7. It's good to add some business intelligence to the talking points, sticking to the market being presented but proving the depth of knowledge of that market by highlighting possible headwinds and tailwinds and how they would react to an inverted yield curve. Also, one should address other issues: If there is an on-going structural shift in demand for homes (are more millennial showing interest in that market); what the NAR's prediction of the summer inventory there is; if the inventory of affordable homes on the rise there; and how any expected change to the FHA rules would help first-time homebuyers in general, etc. 

8. One must try to control the conversation by sticking to what one is presenting, rather than what one does not have. For example, out of the ten variables, if only three are used (e.g., Sale Price, Sale Date, and Bldg SF), one should not start a conversation about the other important variables – Lot size, Age, Bldg Characteristics, and Location – that had to be left out ('If I had 30 more minutes' would be unnecessary). If that question comes up, one must answer it intelligently and truthfully, emphasizing, of course, the added utility of the three variables being used.

9. Now, let's assume that one has managed to complete the first cycle (as indicated above) in 20 minutes. In that case, one must go back to SAS and crunch the sales analysis by the sub-markets (Remember: Location! Location! Location!). In other words, one must understand how to walk down on the analysis curve. 

Of course, it's good to have these printouts handy. Just remember, one complete solution is always better than the more aspiring one but 95% complete.

-Sid Som, MBA, MIM
homequant@gmail.com

Wednesday, September 16, 2020

Post Pandemic, Major Assessment Jurisdictions should consider Hiring AI Engineers (rather than Traditional Modelers)

"AI engineers don't write code to build scalable data pipelines like a data engineer...instead, they understand how to extract data efficiently from a variety of sources, build and test their own machine learning models, and deploy those models using either embedded code or API calls to create AI-infused applications."

Conversely, the existing mass appraisal (CAMA) regression models are not intelligent enough as they are highly modeler dependent (subjective). Thus, given the same sales dataset, five modelers may come up with five different models with very different results. Of course, the biggest failure is the Sales GIS -- generally developed off the current market attributes -- so it's representativeness relative to the population (which is more or less static) is difficult to establish and justify.

[FYI -- That is why, in my AVM books, I propose the use of fixed neighborhoods as they are not sales-dependent and are population-derived, rather than the Sales GIS, which is totally sales-dependent and is not necessarily representative of the population the model is applied to.]

AI engineers do not use any data-variable modeling. Their data extraction process is ingenious, leading to brilliant machine learning models. In a mass appraisal environment, they will precisely identify and demonstrate where the sales datasets and populations are at variance. Whereas, the mass appraisal modelers will remove them from the model as outliers, creating unexplainable gaps and significant fault lines they won't even know.

Alternatively, in a traditional CAMA environment, it's all sample-based, so the results are, at best, bell-curved with the customary 68% efficiency. The error-based CAMA regression models fail to test the solution; for example, is a model Coefficient of Dispersion (COD) of 8 better than a COD of 10? The COD of 8 could represent a post-optimal solution, whereas the COD of 10 could correctly represent an optimal solution. But in a CAMA environment, the COD of 8 would be universally preferred (In fact, several years ago, I presented a paper along these lines at a national conference, raising some serious questions). 

The mass appraisal industry is too antiquated, using the 30-year old regression modeling and mostly Sales GIS. That is why it is high time that the major mass appraisal jurisdictions start hiring some AI engineers, proving that the industry needs to look ahead. 

Granted, given the paucity of AI engineers, it will not be easy to hire AI engineers, but the agencies should widen the search and try. Given how the union-heavy civil service system works, they should also remember what Steve Jobs said, "It does not make sense to hire smart people and then tell them what to do. We hire smart people to tell us what to do." In other words, they must be given the necessary flexibility and autonomy to develop forward-looking solutions, without being bogged down to backward-bending maintenance. 

Of course, to make the modeling environment more efficient and solution-oriented, these agencies should also hire STEM graduates instead of traditional business and humanities graduates who lack advanced quantitative training and knowledge and make very poor modelers. Since a sizable percentage of municipal hires are non-civil servants, these folks could easily qualify in that segment, citing an urgent need for high-level quantitative talent – just the way the major US companies hire skilled foreign nationals under the annual H1-B visa quotas. 

The combination of STEM and AI could be the nirvana for these major jurisdictions.  

Similarly, in the futuristic consumer environment (e.g., free home valuation apps and online sites), the AI-based solutions would, optionally, ask the first-time users to take a short tutorial. As the user interacts with the tutorial, the machine learning models will extract and store the data (by reading the user's responses) and fine-tune the model for each user. When the user returns to value a subject, the stored model will populate the comps as soon as the issue is defined so that the ten-minute exercise would be reduced to fifteen seconds – and customized.

-Sid Som
homequant@gmail.com


Friday, September 11, 2020

Want Fair and Equitable Assessment? Consider Electing New Generation of Tech-Savvy Assessors!

Why?

Here are the reasons: 

1. Independently Elected Assessors will be Free from City Mayor or County Executive's Political Agenda -- Until very recently, Tax Rolls used to be generally regressive, favoring the rich at the expense of the middle class, meaning the middle class used to subsidize - often heavily - the wealthy homeowners. Now with the introduction of the SALT Cap, the wealthy homeowners are not happy campers anymore. Therefore, it is high time to return to the independently Elected Assessors replacing the hand-picked political appointments. Municipal political positions could stoop to shallow levels. Since "home" comprises the single most significant financial investment for the vast majority of taxpayers, they need to regain control of property assessment and, in turn, their property taxes. Electing Mayor/County Executive must be mutually exclusive (de-coupling Administration and Assessment), so they are neither armed with the power of assessment nor bogged down by its never-ending complexities.    

2. Elected Assessors will be Enticed to Undertake Forward-looking Experiments rather than just Vaguely Meeting some Antiquated Industry Guidelines -- Unfortunately, as long as some antiquated industry guidelines are met, the Roll gets published, with the Review Board picking up the fall-outs from there. This feedback loop (feeding off each other) perennially continues while the taxpayers are left in the lurch. Meanwhile, the current Administration continues to point the finger at the former Mayor/County Executive of slashing the human resources budget. A visionary and gutsy Elected Assessor would look at such scenarios positively and explore forward-looking solutions instead of more government employees. Taxpayers need a real solution to this age-old challenge. Independent and innovative leadership could be the step in the right direction. The old-age experiment of more government employees has not worked.

3. The New World of Artificial Intelligence could provide the Solution the Assessment Industry has been Waiting for -- if the new generation Elected Assessors are keen on having this age-old challenge solved, they should seriously consider the world of AI solutions. Therefore, an Elected Assessor needs to be someone with a background in solving complex problems, perhaps without assessing or Valuing whatsoever. A brilliant problem-solving mind will quickly realize the need for AI/Robotics technology in getting arms around this challenge. Such reasons would not buy into the idea of the old-fashioned statistical modeling with high "under-the-surface" error rates. For example, an AI-based data collection and data update application will do a far better job than competing human judgment (despite the homogeneous definitions and guidelines).    

4. The New Generation of Elected Assessors will help Change the Existing Guidelines, which have become vastly Antiquated -- The existing valuation modeling (CAMA) guidelines revolve around the multiple regression models. That modeling environment provided an excellent start to the industry, but, over the years, the over-dependence on the modeling stats to meet and exceed the industry guidelines often paved the way for significantly less surgical dissection and analysis of the population data, resulting in sizable appeals, especially in urban jurisdictions with a high degree of heterogeneous housing stock, thus calling into question the effectiveness of the overall Roll. A new generation of Elected Assessors with advanced quantitative and problem-solving backgrounds would be needed to introduce rules leading to meaningful guidelines.

5. Elected Assessors would be more Amenable to FOIA/FOIL Requests on Automated Valuation Models and Formulas -- In the event of an error-filled Roll, the Assessors who are part of the ruling Administration would be less than forthcoming to disclose their models and formulas, whether internally developed or by their favorite consultants, to the public. The new generation of Elected Assessors would be more than happy to make them public, emphatically citing the progress that has been made and highlighting the pipeline that is being worked on, even if it involves taking on some short-term pain to achieve significant long-term gain (i.e., to solve the age-old challenge). They might even encourage a series of brain-storming debates with local experts to factor their inputs into the process. Playing hide-and-seek game with the taxpayers is antagonistic to the fortitude of public service. 

6. Elected Assessors will Meaningfully Evaluate Cost Benefits of Outsourcing certain Functions and Services vis-a-vis Conventional New Hires, Consultants and Vendors -- The new generation of Elected Assessors will bring a wealth of expertise in making sound business decisions, including BPOs, KPOs, etc. Their choices will conform to the election manifesto that put them in the office in the first place given their independent status (off of the ruling Administration). The knowledge of the make-shift internal modelers (from assessing/data collection to quantitative modeling) and the expertise the industry's leading consultants and vendors bring are perhaps backward-bending, forcing them to consider other more forward-looking alternatives like hiring AI Engineers, outsourcing to AI firms, retaining Public Relations firms for outreach, etc. The alternative is the status quo, with wishful thinking (will fix the next Roll). 

7. Elected Assessor's Vision will help many Veterans to leave yielding place to the New -- Tennyson's immortal poem may ring a bell again, "The old order changeth yielding place to new And God fulfills himself in many ways Lest one good custom should corrupt the world." As the new vision anchors, many veterans unwilling or unable to keep pace with the changes might retire, freeing up significant budget dollars for the modern workforce, technology, and futuristic services. Since the veterans are some of the highest-paid employees, the savings from their departure could be significant. Perhaps, the new generation of Assessors' election might finally work as the catalyst or the leading indicator of change for the other interacting agencies as well, proving how forward-looking vision backed by disruptive business models could bring about system-wide changes.   

8. Elected Assessors might take a fresh look at the existing Data Warehousing and Modeling -- If only 5-10% of the data undergoes annual changes, the question one should ask: Is it worthwhile to spend millions of taxpayer dollars on warehousing the static data or is it prudent to switch to a significantly scaled-down data warehouse? A straight-forward AI model based on just size and location could be more effective than a regression-based Juggernaut requiring a whole host of modelers applying personal judgment to improve model stats. Therefore, future data warehousing should be tied to future valuation modeling, eliminating the need for any wasteful data. If a house has two full baths, who cares about keeping track of the number of bath fixtures? Just an extravagant piece of data and costly warehousing! Therefore, the new generation of Elected Assessors would be expected to cut to the chase as to the future data requirements, saving a ton on wasteful data collection, inspection, and warehousing.   

9. Appeals Review Assessor must also be Elected -- Almost all major jurisdictions have Appeals Review agencies (under various names) headed by Commissioners, etc., usually hand-picked by the ruling Administration. While this separation (generally by the Charter) might be useful, their total independence is more critical to effectively serving the taxpayers. Therefore, to make the system politics-free, taxpayers must have "Elected" Review Assessors as well. While electing the Review Assessor, taxpayers should zero in on the candidates with significant Quality Control/Assurance expertise. Of course, independence comes with a financial cost, meaning an independent apparatus has to be maintained. Since the review window is short-lived (seasonal), employees could be leased to avoid employing them during the off-peak season, incurring unnecessary legacy costs.    

There is no certainty that an Elected Assessor or Review Assessor would always pan out; then again, the power remains with the people.

Of course, the only permanent solution to this perennial problem is to gradually replace property taxes with middle-class friendly progressive consumption taxes (see the link below). 

-Sid Som, MBA, MIM
homequant@gmil.com             

Post Pandemic, Replace Property Taxes with Middle-Class friendly Progressive Consumption Taxes         


Wednesday, January 15, 2020

How to Analyze and Present a Complex Dataset – in 60 Minutes

   ** For New Graduates/Analysts **

We talked about analyzing and presenting a large and complex dataset in 30 minutes in the prior blog post. Would one handle it differently if one had 60 minutes? Here is one approach one might like to consider:

1. While starting out, many young folks tend to underestimate themselves. The very fact that one has been tasked with this critical presentation speaks volumes, so one must learn to take full advantage of this visibility in narrowing the (internal) competition down. These meetings are often frequented by other department heads and high-level client representatives, leading to significant loss of time in unrelated (business) discussions. The best way to prepare for such contingencies is to split the presentation into a two-phase solution where phase-1 leads seamlessly to phase-2. 

2. In a business environment, it's never a good idea to start with a complicated stat/econ model; instead, one must start a bit slow but use one's analytical acumen and presentation skill to gradually force people to converge on the same page, retaining maximum control over the presentation in terms of both time and theme). Therefore, the phase-1 solution should be the same as the full 30-minute solution we detailed in a prior blog post (including the sub-market analysis). Even if the meeting leads to unrelated business chit-chat, off and on, the presenter will still be able to squeeze in the phase-1 solution, thus offering at least a baseline solution. Alternatively, if one has an all-encompassing solution, one could end up offering virtually nothing. 

3. Now that the phase-1 presentation, establishing a meaningful baseline is over, one should be ready to transition to the higher-up phase-2 solution. In other words, it's time to show off one's modeling knowledge. The phase-1 presentation comprised a baseline Champ-Challenger analysis, where the Champ was the Monthly Median Sale Price, and the Challenger was the Monthly Median SP/SF. The presenter used the "Median" to avoid having to clean up the dataset for significant outliers. Here is the caveat of sales analysis though: Sales, individually, are mostly judgment calls; for example, someone bent on buying a pink house would overpay; an investor would underpay by luring a seller with a cash offer, etc. In the middle (middle 68% of the bell curve), the so-called informed buyers would use five comps, usually hand-picked by the salespeople, to value the subjects – not an exact science either.  

4. Now, let's envision where the presenter would be at this stage – 30 minutes on hand and brimming with confidence. But it's not enough time to try to develop and present an accurate multi-stage, multi-cycle AVM. So, it's good to settle for a straight-forward regression-based modeling solution, allowing time for a few new slides to the original presentation. Ideally, the model should be built as one log equation with a limited number of variables (though covering all three major categories). The variables one might like to choose are: Living Area, Age, Bldg Style, Grade, Condition, and School/Assessing District, avoiding the 2nd tier variables (e.g., Garage SF, View, Site Elevation, etc.).

5. One should use Time Adjusted Sale Price (ASP) as the dependent variable in the Regression model, explaining the connection between the presentations (meaning phase-1 and 2) so the audience (including the big bosses like the SVP, EVP, etc.) understands that the two phases are not mutually exclusive, rather one is the stepping stone to the other. At this point, the presenter could face this question "Why did you split it up into two?" The answer must be short and truthful: "It's a time-based contingency plan."

6. At this point, the presenter must keep the regression output handy without inserting it into the main presentation, though, considering it is a log model (the audience may not relate to the log parameter estimates). If the issue comes up, the presenter should talk about the three critical aspects of the model: (a) the variable selection (how all of the three categories were represented), (b) the most vital variables as judged by the model (walking down on the t-stat and p-value), and (c) overall accuracy of the model (zeroing on the primary stats like r-squared, f-statistics, confidence, etc.).   

7. The presenter must explain the model results in three simple steps: (a) Value Step: ASPs vs. Regression values, showing the entire percentile curve, 1st to 99th percentile rather than the median values only, and also pointing out the inherent smoothness of the Regression values vis-a-vis the ASPs; (b) Regression Step: How some arms-length sales could be somewhat irrational on both ends of the curve (<=5th and >=95th) and why the standard deviation of the Regression values was so much lower than ASP'; and (c) Ratio Step: Stats on the Regression Ratio (Regression Value to ASP) as it's easier to explain the Regression Ratios than the natural numbers so spending more time on the ratios would make the presentation more effective.   

8. The presenter should explain the outlier ranges -- the ratios below the 5th and above the 95th percentile, or below 70 and above 143. Considering this is the outlier-free output, it's good to display Std Dev, COV, COD, etc. The outlier-free stats would be significantly better than the prior (with outliers) ones. Another common outlier question is: "Why no waterfront in your model?" The answer is simple: Generally, waterfront parcels comprise less than 5% of the population, hence challenging to test representativeness. (In an actual AVM, if sold waterfront parcels properly represent the waterfront population, it could be tried in the model, as long as it clears the multi-collinearity test as well).  

9. Last but least, one must be prepared to face an obvious question: "What is the point of developing this model?" Here is the answer: "A sale price is more than a handful of top-line comps. It comprises an array of important variables like size, age, land, building characteristics, fixed and micro-locations, etc. so only a multivariate model can do justice to sell prices by properly capturing and representing all of these variables. The output from this Regression model is the statistically significant market replica of the sales population. Moreover, this model can be applied to the unsold population to generate significant market values. Simply put, this Regression model is an econometric market solution. Granted, the unsold population could be comp'd, but that's a very time-consuming and subjective process."

-Sid Som
homequant@gmail.com

How to Analyze and Present a Complex Dataset – in 30 Minutes

** For New Graduates/Analysts **

Often, with minimal time on hand – say 30 minutes – to summarize and present a relatively large and complex home sales dataset, comprising 18 months of data, with 30K rows, and ten variables, here is one approach worth considering:

1. Given the limited time, instead of trying to crunch the data in a spreadsheet, it's better to one's your favorite statistical software like SAS, SPSS, etc. What SAS will do in four short statements (Proc Means, Var, Class, and Output), and a matter of minutes, will need much longer to accomplish the same in spreadsheets. When one is starting out, it's good to take full advantage of these types of highly visible, often gratifying challenges to narrow the potential competition down.

2. It's good to have a realistic game plan. Instead of shooting for an array of parameters, it's better to start with the most significant one, i.e., Monthly Median Sale Price (and the normalized Sale Price per SF). Since the median is not prone to outliers, the dataset doesn't have to be edited for outliers, saving a significant amount of time.  

3. Now that the monthly median prices are there, one should be ready to create graphs for the presentation. While one graph depicting both prices (Y1 and Y2) against months (X-axis) may be created, it's prudent to keep them separated for ease of presentation. 

4. Since basic graphing is more straightforward in Excel (in fairness to the remaining time), it's better to transfer the output from SAS to Excel, ensuring that the graphs are adequately annotated and dressed up with the axis titles, legends, gridlines, etc. One must also remember that just doing things the right is not good enough, one must learn to present things elegantly as well. 

5. Since so much of the data have been summarized and rolled up behind one or two graphs, one must make sure they not only tell the overall story but also convey enough business intelligence to make the presentation look like a well-thought-out business solution in front of the attending EVP, SVP, etc. In the presence of clients, it enhances the bosses' image as well. So, it's smart to add trendlines alongside the data trend, selecting the primary trendline by eyeballing the data trend (linear, logarithmic, polynomial, etc.). Adding a 2 to 3-month moving average (depending on the time series) trendline to iron out any monthly aberrations could enhance the presentation.

6. It's also smart to keep the reporting verbiage clear and concise, explaining the makeup of the dataset, methodology including monthly medians, and how the normalized prices add value and help validate the primary. It's also important to explain the use of the trendline and its statistical significance and the other statistical measures like r-squared, slopes, etc. one might display on the graphs (avoiding the printing of equations on the graphs). 

7. It's good to add some business intelligence to the talking points, sticking to the market being presented but proving the depth of knowledge of that market by highlighting possible headwinds and tailwinds and how they would react to an inverted yield curve. Also, one should address other issues: If there is an on-going structural shift in demand for homes (are more millennial showing interest in that market); what the NAR's prediction of the summer inventory there is; if the inventory of affordable homes on the rise there; and how any expected change to the FHA rules would help first-time homebuyers in general, etc. 

8. One must try to control the conversation by sticking to what one is presenting, rather than what one does not have. For example, out of the ten variables, if only three are used (e.g., Sale Price, Sale Date, and Bldg SF), one should not start a conversation about the other important variables – Lot size, Age, Bldg Characteristics, and Location – that had to be left out ('If I had 30 more minutes' would be unnecessary). If that question comes up, one must answer it intelligently and truthfully, emphasizing, of course, the added utility of the three variables being used.

9. Now, let's assume that one has managed to complete the first cycle (as indicated above) in 20 minutes. In that case, one must go back to SAS and crunch the sales analysis by the sub-markets (Remember: Location! Location! Location!). In other words, one must understand how to walk down on the analysis curve. 

Of course, it's good to have these printouts handy. Just remember, one complete solution is always better than the more aspiring one but 95% complete.

-Sid Som, MBA, MIM
homequant@gmail.com